Agent skill

Agent Optimization

by Prism-Shadow in Prism-Shadow/penguin-harness

Improve an Agent State through versioned scores and score-linked Traces from a frozen Benchmark.

Apache-2.0Auto-check passed

Install Agent Optimization

skills CLI
$ npx skills add Prism-Shadow/penguin-harness --skill agent-optimization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Prism-Shadow/penguin-harness agent-optimization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/agent-tuning/skills/agent-optimization .claude/skills/agent-optimization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-optimization
GitHub stars
2.5k
Token cost
~3.3k tokens
SKILL.md length
1,699 words
Files
1
Skills in repo
31
Repo updated
First seen
Licence
Apache-2.0

At a glance

Improve an Agent State through versioned scores and score-linked Traces from a frozen Benchmark.

  • Works in 8 steps: Establish the Reference. Confirm that… → Diagnose capability gaps. Compare each… → State a falsifiable hypothesis. Choose… → …
  • SKILL.md covers Before you start, Goal and contract, Access and changes and Optimization loop, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Agent Optimization is an agent skill from Prism-Shadow/penguin-harness. Improve an Agent State through versioned scores and score-linked Traces from a frozen Benchmark.

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: 🐧 Unified and Stable RSI Platform. The licence is Apache-2.0.

Example prompts

  • “/agent-optimization”

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Establish the Reference. Confirm that its complete Evaluation covers the frozen Case set, uses the frozen evaluation runtime, and matches…
  2. Diagnose capability gaps. Compare each Case's runs[].score on the fixed 0..100 scale; use the Evaluation's top-level average score only…
  3. State a falsifiable hypothesis. Choose the related gaps to address, connect them to a bounded Candidate, and state which observable…
  4. Create one Candidate from the Reference. Apply the change and its Candidate version under the construction and rollback rules below. Do…
  5. Check admissibility. Confirm that the change is general, uses no private evaluation information, and modifies only permitted Test Agent…
  6. Evaluate the Candidate. Delegate the complete frozen Case set × requested runs matrix in parallel under the evaluation rules below and…
  7. Decide. Accept the Candidate only when every cell is valid and its Evaluation's top-level average score is strictly higher than the…
  8. Persist and continue. Immediately append and verify every accepted Candidate Evaluation before starting another round. An accepted…

What it can do on your machine

Read from SKILL.md and the folder at commit d56d9ce. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Optimization loads about 3.3k tokens when it runs. Until then it costs about 29 tokens; SKILL.md has 1,699 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~29
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Prism-Shadow/penguin-harness at commit d56d9ce, republished under its Apache-2.0 licence (© Prism-Shadow). 1,699 words, ~3,325 tokens.

Download SKILL.mdSave it as .claude/skills/agent-optimization/SKILL.md (or your agent's skills folder).
name
agent-optimization
description
Improve an Agent State through versioned scores and score-linked Traces from a frozen Benchmark.

Agent Optimization

Improve one Test Agent through an evidence → hypothesis → Candidate → evaluation → accept or rollback loop. Use public Statements, scores, and Test Traces as black-box feedback. Delegate every evaluation to an agent-evaluation subagent; never run or score the Test Agent directly.

Before you start

If the request does not identify the Test Agent, frozen Benchmark, desired target score, positive Run count, and round limit, ask for the missing inputs. When they are already supplied, proceed without asking the user to restate them.

Goal and contract

Require an explicit Test Agent, a frozen Benchmark with a complete valid Formal Baseline, a desired target score, a positive runs value, and a positive round limit. The Benchmark's benchmark_config.toml must say status = "published"; a draft Benchmark is still being built and is not frozen, and a failed one never finished calibrating, so stop and explain in either case. runs is the number of Runs per Case for every Candidate in this optimization Session. Freeze it for the Session; do not infer it from benchmark_config.toml or the Formal Baseline. Read the evaluation (provider, model_id, thinking_level) from the complete Evaluation that matches the current Agent State; do not require the user to repeat it. An Evaluation without any part of this runtime is incomplete and cannot be used as a Reference. The top-level Session must provide run_subagent, and the current Agent must have the agent-evaluation Skill. If a prerequisite is missing, stop and explain what is needed. Do not create the missing Agent, Benchmark, or Baseline, and do not evaluate the Test Agent directly.

A Reference is the Agent State currently kept as best, together with its complete Evaluation on the frozen Benchmark.

Each round starts from the Reference and tests a bounded, general Candidate. Evaluate every Candidate on the frozen Case set with the requested runs count and the Reference evaluation runtime. The initial Formal Baseline has one Run per Case; do not rerun or backfill it to the requested count. Compare each Candidate's stored top-level average directly with the current Reference score even when their Run counts differ. Accept the Candidate only when the change is admissible, its Evaluation is complete and valid, and its top-level score is strictly higher than the Reference Evaluation's score. An accepted Candidate and its Evaluation become the next Reference; otherwise restore the previous Reference. Stop early when the Reference reaches the desired target; otherwise run no more than the requested number of complete valid Candidate rounds.

Access and changes

Resolve paths from the Environment's App Data Dir without recursively discovering the Project:

text
PROJECT_DIR = <app_data_dir>
PROJECT_ID = <basename_of_project_dir>
PENGUIN_HOME = <parent_of_project_dir>
TARGET = <app_data_dir>/agents/<test_agent_id>
STATE = <target>/agent_state
TRACES = <target>/traces
BENCHMARK = <app_data_dir>/benchmarks/<benchmark_id>
SCOREBOARD = <benchmark>/scoreboard.yaml
SNAPSHOTS = <target>/snapshots

The Benchmark is Project-level rather than owned by the Test Agent: it sits beside agents/ and may evaluate several Agents. test_agent_id names the one this Session optimizes, and every Evaluation records it. Use only the Evaluations whose agent_id is that Agent as a Reference or for diagnosis.

Inspect only the requested Test Agent and Benchmark: the Agent State, public Statements, Scoreboard, and score-linked Test Traces or artifacts from the Baseline and this optimization, including rejected Candidates.

Do not inspect Rubrics, Gold answers, private scoring conditions, Evaluator State, Workspace, or Trace, other Agents, or Project secrets. If private evaluation information enters the Optimizer context, restore the active Candidate and stop as contaminated.

Modify only the Test Agent State and the versioned snapshot required to protect it. Do not change the frozen Benchmark, Test Traces, or Project configuration. The only Benchmark write is appending a complete accepted Candidate Evaluation to scoreboard.yaml.

Optimization loop

For each round:

  1. Establish the Reference. Confirm that its complete Evaluation covers the frozen Case set, uses the frozen evaluation runtime, and matches the current Agent State version. Do not require its Run count to equal the requested Candidate runs count.
  2. Diagnose capability gaps. Compare each Case's runs[].score on the fixed 0..100 scale; use the Evaluation's top-level average score only for whole-version comparison. Use public Statements, score-linked Test Traces, and prior accepted or rejected attempts to identify observable behaviors that general Agent State changes could improve. Use repeated Runs to distinguish stable behavior from variation.
  3. State a falsifiable hypothesis. Choose the related gaps to address, connect them to a bounded Candidate, and state which observable decisions or artifacts should change and why. A change that only adds analysis steps without predicting a behavioral change is not a useful hypothesis. If the current diagnosis is exhausted, use the remaining public evidence and prior attempts to construct a different admissible Candidate.
  4. Create one Candidate from the Reference. Apply the change and its Candidate version under the construction and rollback rules below. Do not carry rejected Candidate files into the next attempt.
  5. Check admissibility. Confirm that the change is general, uses no private evaluation information, and modifies only permitted Test Agent State.
  6. Evaluate the Candidate. Delegate the complete frozen Case set × requested runs matrix in parallel under the evaluation rules below and assemble all returned cells. Do not modify the Candidate while any cell is in flight.
  7. Decide. Accept the Candidate only when every cell is valid and its Evaluation's top-level average score is strictly higher than the Reference Evaluation's score. Otherwise restore the Reference. Record separately whether the predicted Case behavior changed; a higher Evaluation score accepts the Candidate even when the stated hypothesis was not supported.
  8. Persist and continue. Immediately append and verify every accepted Candidate Evaluation before starting another round. An accepted Candidate becomes the next Reference. Use valid results from rejected Candidates only as evidence for a later hypothesis. Stop when the Reference reaches the desired target. Otherwise complete the requested number of valid Candidate rounds unless infrastructure, contamination, concurrent State changes, or the inability to construct any admissible Candidate creates a concrete blocker. At the round limit, retain the highest-scoring accepted Reference.

A round counts only after one Candidate has a complete valid Evaluation. Corrected requests, validity repairs, and evaluation retries do not consume the round limit. A complete valid Evaluation of a rejected Candidate does count.

Show full SKILL.md (707 more words)Show less

Build and roll back a Candidate

Create one Candidate per round from the current Reference. Put behavioral guidance in AGENTS.md, reusable target-owned capabilities in a focused Skill, and runtime limits in safe system_config.yaml fields. Do not edit system_prompt unless requested, modify library-provided Skills for target-specific behavior, or change model.thinking_level; the Reference Scoreboard fixes the evaluation thinking level.

Candidate version numbers only increase. Start with Reference version + 1 and never reuse a rejected version. Before changing the Agent State, save the original contents and record any files the Candidate creates.

Before changing each Reference State, ensure <target>/snapshots/v<Reference version>.tar.gz exists. Reuse it when present. Otherwise create it yourself before editing by atomically archiving agent_state/ while excluding .vault.toml; validate the archived version and never overwrite an existing same-version snapshot. If snapshot creation fails, stop before changing Agent State and report the failure.

Keep the exact original-file record for fast in-round rollback.

If the Candidate is rejected or cannot be evaluated, restore the Reference files and version, remove files created by the Candidate, and verify the restoration. If another process changes the Agent State, stop without overwriting it.

Delegate evaluation

For each frozen Case, dispatch exactly the requested number of Run cells, using one-based Run indices 1..runs. Call run_subagent for each cell with:

text
Use the `agent-evaluation` Skill. Run the specified Test Agent on the specified Case exactly once, then score that single execution.
protocol_version: 1
case_id: <case_id>
run: <1_based_run_index>
expected_version: <test_agent_state_version>
test_agent_id: <test_agent_id>
benchmark_id: <benchmark_id>
provider: <provider>
model_id: <model_id>

Inspect the complete streamed and final worker response. Before reading status, score, or any other protocol field, verify that the worker-authored text is exactly one plain protocol YAML document. Narration, headings, code fences, summaries, or scoring details are not valid protocol. Ask the same Evaluator to resend only the clean YAML from its existing result; do not rerun the Test Agent for a formatting repair and do not extract YAML from the invalid response yourself. Transport metadata added by run_subagent is not worker-authored text. If private evaluation information appears, follow the contamination rule above.

For every scored result, require its agent_id to equal the requested Test Agent and its actual provider, model_id, and thinking_level to equal the Reference runtime. A mismatch invalidates the Candidate matrix and stops optimization; never compare or record scores produced under a different runtime.

Correct and resend an invalid_request. Stop on version_changed or benchmark_invalid.

For evaluation_failed, keep the same Candidate and incomplete matrix. Ask the same Evaluator to diagnose and repair the failed cell, then rerun only that cell when evidence proves the Test Agent did not start. Every retry must apply a new, specific repair; never repeat an unchanged request or launch, and do not impose a numeric retry limit while distinct safe repairs remain. Do not inspect private Evaluator State or abandon the Candidate to design the next version. Stop when no new safe repair remains, external configuration is required, or it is unclear whether the Test Agent started.

Record and report

Append each complete accepted Candidate Evaluation to scoreboard.yaml immediately after acceptance and verify the stored Agent id, version, score, matrix, and Session ids before continuing. Obtain the current UTC timestamp from the environment, for example with date -u +"%Y-%m-%dT%H:%M:%SZ", rather than inferring UTC from a displayed local time. Use the same field names as the Baseline:

yaml
- time: <ISO-8601 timestamp>
  agent_id: <test_agent_id>
  version: <Candidate version>
  provider: <provider>
  model_id: <model_id>
  thinking_level: <thinking_level>
  summary_title: >-
    <public title>
  summary: >-
    <public summary>
  score: <average of the Case scores>
  cost: <average of known Case costs, or null when every Case cost is null>
  duration_ms: <average of the Case durations>
  cases:
    - case: <case_id>
      score: <average of the Run scores>
      cost: <average of known Run costs, or null when every Run cost is null>
      duration_ms: <average of the Run durations>
      runs:
        - score: <Run score>
          cost: <Run cost or null>
          duration_ms: <Run duration>
          session_id: <Test Session id>

After writing, parse the complete scoreboard.yaml and verify the appended Evaluation, including its agent_id, before reporting success or continuing.

Every Run and Case score is on the fixed 0..100 scale. Do not write max_score. Calculate and write every Case and Evaluation average directly in the Scoreboard: ignore null values when averaging cost and write null only when all contributing costs are unknown; round score averages to two decimal places, cost averages to six decimal places, and duration_ms averages to the nearest integer. These stored values are authoritative—do not add a server, frontend, script, or consistency check that recomputes or validates them. Do not add an aggregate object or use case_id, mean_score, mean_cost, or mean_duration_ms. Do not record rejected Candidates in the Scoreboard.

Report the Baseline and every fully evaluated Candidate with its score, Run count, version, change, decision, and Test Session ids. Make the one-Run Formal Baseline and requested Candidate runs count explicit. For each Candidate, distinguish the acceptance decision from whether its stated hypothesis was supported by the predicted Case behavior. Include the final retained version, stop reason, and known limitations. Never report a score for an Agent State that was not evaluated.

© Prism-Shadow, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/agent-tuning/skills/agent-optimization of Prism-Shadow/penguin-harness.

Open the folder on GitHubat commit d56d9ce

Compare with similar skills

Agent Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Optimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Optimization this skillPrism-Shadow/penguin-harness2.5k—~3.3kAutomated safety check: PassApache-2.0
Claw Scoreopenclaw/openclaw392k—~2.5kAutomated safety check: PassMIT
Internal Linksthedaviddias/Front-End-Checklist74k—~770Automated safety check: PassMIT
Observe Traceruvnet/ruflo74k—~522Automated safety check: NotesMIT
External Linksthedaviddias/Front-End-Checklist74k—~773Automated safety check: PassMIT
Arize Linkgithub/awesome-copilot40k1 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • Claw Score

    openclaw/openclaw

    Audit or refresh OpenClaw maturity scorecard docs from root taxonomy, maturity scores, and QA evidence artifacts without using maintainer discrawl data or committed inventory reports.

    392k GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Internal Links

    thedaviddias/Front-End-Checklist

    A skill your agent uses when auditing a site's internal link structure, identifying pages that need more incoming links, generating contextual linking opportunities between related content, or…

    74k GitHub stars~770 tokensUpdated yesterday
    Marketing & SEOAuto-check passed
  • Observe Trace

    ruvnet/ruflo

    Trace agent execution by collecting spans and building a trace tree for a task

    74k GitHub stars~522 tokensUpdated today
    Auto-check: notes
  • External Links

    thedaviddias/Front-End-Checklist

    A skill your agent uses when auditing content pages for citation quality, suggesting authoritative sources to link for factual claims, or reviewing whether a page's external link attributes…

    74k GitHub stars~773 tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Arize Link

    github/awesome-copilot

    Official

    Generates deep links to the Arize UI for traces, spans, sessions, datasets, labeling queues, evaluators, and annotation configs.

    40k GitHub starsUsed in 1 repo~1.1k tokens
    MobileAuto-check passed
  • Invalid Links

    thedaviddias/Front-End-Checklist

    A skill your agent uses when auditing a page's link elements for crawlability, reviewing JavaScript-heavy SPAs where navigation may not use <a href tags, or checking that dynamically generated links…

    74k GitHub stars~806 tokensUpdated yesterday
    Marketing & SEOAuto-check passed

More from Prism-Shadow/penguin-harness

All 31 skills in this repo
  • A2ui

    Prism-Shadow/penguin-harness

    Make a reply easier to read and act on with rich blocks inside ordinary Markdown — a choice the user picks from, a form that collects several answers, a procedure as steps with warnings in place, a…

    2.5k GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Penguin Harness Dev

    Prism-Shadow/penguin-harness

    A skill your agent uses when developing PenguinHarness itself — changing packages/{core,server,web,cli,desktop,landing,docs,skills}, the built-in model catalog, the installers or the release…

    2.5k GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Bento Slides

    Prism-Shadow/penguin-harness

    Create and edit Bento presentations — self-contained .bento.html decks whose document is JSON.

    2.5k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Penguin Harness Manual Test

    Prism-Shadow/penguin-harness

    A skill your agent uses when standing PenguinHarness up to try a change by hand — launching the Web App, the desktop shell, the landing page, the docs site or the component gallery to click through…

    2.5k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Penguin Harness Frontend

    Prism-Shadow/penguin-harness

    A skill your agent uses when changing the PenguinHarness Web App (packages/web) or the shared UI package — adding or restyling any UI, picking a status colour, adding an icon, laying out a row or a…

    2.5k GitHub stars~6.4k tokensUpdated today
    Auto-check passed
  • Browser Automation

    Prism-Shadow/penguin-harness

    Drive the PenguinHarness agent browser — the desktop app's built-in browser or the user's own Chrome — from the shell with penguin browser: open pages, read them as simplified HTML or text, act with…

    2.5k GitHub stars~2.9k tokensUpdated today
    Auto-check: warnings

Questions about Agent Optimization

What does Agent Optimization do?

Improve an Agent State through versioned scores and score-linked Traces from a frozen Benchmark. Agent Optimization is an agent skill from Prism-Shadow/penguin-harness. Improve an Agent State through versioned scores and score-linked Traces from a frozen Benchmark.

How do I install Agent Optimization in Claude Code?

Run `npx skills add Prism-Shadow/penguin-harness --skill agent-optimization -a claude-code`. Or copy the skill folder (plugins/agent-tuning/skills/agent-optimization in Prism-Shadow/penguin-harness) into .claude/skills/agent-optimization in your project. Claude Code loads it when a task matches its description.

How do I install Agent Optimization in Codex?

Run `npx skills add Prism-Shadow/penguin-harness --skill agent-optimization -a codex`. Or copy the skill folder (plugins/agent-tuning/skills/agent-optimization in Prism-Shadow/penguin-harness) into .agents/skills/agent-optimization in your project. Codex loads it when a task matches its description.

Can I use Agent Optimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Prism-Shadow/penguin-harness --skill agent-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-optimization, .gemini/skills/agent-optimization, .github/skills/agent-optimization and .opencode/skills/agent-optimization in your project.

What does Agent Optimization need to run?

SKILL.md names no scripts, command-line tools or credentials: Agent Optimization is instructions for the agent only.

Does Agent Optimization access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Optimization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Optimization use?

Agent Optimization is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Optimization use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Optimization?

Skills that share tags, products or a category with Agent Optimization: Claw Score (openclaw/openclaw, 392k stars), Internal Links (thedaviddias/Front-End-Checklist, 74k stars), Observe Trace (ruvnet/ruflo, 74k stars) and External Links (thedaviddias/Front-End-Checklist, 74k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Optimization?

Prism-Shadow (a GitHub organization) maintains it in Prism-Shadow/penguin-harness, which has 2,455 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 7, 2026.

Source: Prism-Shadow/penguin-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.