Agent skill

Evaluate Presets

by mikeyobrien in mikeyobrien/ralph-orchestrator

A skill your agent uses when testing Ralph's hat collection presets, validating preset configurations, or auditing the preset library for bugs and UX issues.

MITAuto-check passed

Install Evaluate Presets

skills CLI
$ npx skills add mikeyobrien/ralph-orchestrator --skill evaluate-presets -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mikeyobrien/ralph-orchestrator evaluate-presets --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mikeyobrien/ralph-orchestrator.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/evaluate-presets .claude/skills/evaluate-presets && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluate-presets
GitHub stars
3.2k
Token cost
~1.8k tokens
SKILL.md length
625 words
Files
1
Skills in repo
16
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when testing Ralph's hat collection presets, validating preset configurations, or auditing the preset library for bugs and UX issues.

  • Works in 4 steps: Triage Results → Dispatch Task Creation → Dispatch Implementation → …
  • Testing Ralphs hat collection presets
  • SKILL.md covers Overview, When to Use, Quick Start and Bash Tool Configuration, plus 7 more sections
  • Calls bash, jq and brew

What it does

Evaluate Presets is an agent skill from mikeyobrien/ralph-orchestrator. Use when testing Ralph's hat collection presets, validating preset configurations, or auditing the preset library for bugs and UX issues.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: An improved implementation of the Ralph Wiggum technique for autonomous AI agent orchestration. The licence is MIT.

When your agent uses it

  • Testing Ralphs hat collection presets
  • Validating preset configurations
  • Auditing the preset library for bugs and UX issues

Example prompts

  • “/evaluate-presets”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Triage Results
  2. Dispatch Task Creation
  3. Dispatch Implementation
  4. Re-evaluate

What it can do on your machine

Read from SKILL.md and the folder at commit edc2b32. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bash
    • jq
    • brew

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluate Presets loads about 1.8k tokens when it runs. Until then it costs about 39 tokens; SKILL.md has 625 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~39
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mikeyobrien/ralph-orchestrator at commit edc2b32, republished under its MIT licence (© mikeyobrien). 625 words, ~1,828 tokens.

Download SKILL.mdSave it as .claude/skills/evaluate-presets/SKILL.md (or your agent's skills folder).
name
evaluate-presets
description
Use when testing Ralph's hat collection presets, validating preset configurations, or auditing the preset library for bugs and UX issues.
metadata.internal
true

Evaluate Presets

Overview

Systematically test all hat collection presets using shell scripts. Direct CLI invocation—no meta-orchestration complexity.

When to Use

  • Testing preset configurations after changes
  • Auditing the preset library for quality
  • Validating new presets work correctly
  • After modifying hat routing logic

Quick Start

Evaluate a single preset:

bash
./tools/evaluate-preset.sh tdd-red-green claude

Evaluate all presets:

bash
./tools/evaluate-all-presets.sh claude

Arguments:

  • First arg: preset name (without .yml extension)
  • Second arg: backend (claude or kiro, defaults to claude)

Bash Tool Configuration

IMPORTANT: When invoking these scripts via the Bash tool, use these settings:

  • Single preset evaluation: Use timeout: 600000 (10 minutes max) and run_in_background: true
  • All presets evaluation: Use timeout: 600000 (10 minutes max) and run_in_background: true

Since preset evaluations can run for hours (especially the full suite), always run in background mode and use the TaskOutput tool to check progress periodically.

Example invocation pattern:

Bash tool with:
  command: "./tools/evaluate-preset.sh tdd-red-green claude"
  timeout: 600000
  run_in_background: true

After launching, use TaskOutput with block: false to check status without waiting for completion.

What the Scripts Do

evaluate-preset.sh
  1. Loads test task from tools/preset-test-tasks.yml (if yq available)
  2. Creates merged config with evaluation settings
  3. Runs Ralph with --record-session for metrics capture
  4. Captures output logs, exit codes, and timing
  5. Extracts metrics: iterations, hats activated, events published

Output structure:

.eval/
├── logs/<preset>/<timestamp>/
│   ├── output.log          # Full stdout/stderr
│   ├── session.jsonl       # Recorded session
│   ├── metrics.json        # Extracted metrics
│   ├── environment.json    # Runtime environment
│   └── merged-config.yml   # Config used
└── logs/<preset>/latest -> <timestamp>
evaluate-all-presets.sh

Runs all 12 presets sequentially and generates a summary:

.eval/results/<suite-id>/
├── SUMMARY.md              # Markdown report
├── <preset>.json           # Per-preset metrics
└── latest -> <suite-id>

Presets Under Evaluation

PresetTest Task
tdd-red-greenAdd is_palindrome() function
adversarial-reviewReview user input handler for security
socratic-learningUnderstand HatRegistry
spec-drivenSpecify and implement StringUtils::truncate()
mob-programmingImplement a Stack data structure
scientific-methodDebug failing mock test assertion
code-archaeologyUnderstand history of config.rs
performance-optimizationProfile hat matching
api-designDesign a Cache trait
documentation-firstDocument RateLimiter
incident-responseRespond to "tests failing in CI"
migration-safetyPlan v1 to v2 config migration

Interpreting Results

Exit codes from evaluate-preset.sh:

  • 0 — Success (LOOP_COMPLETE reached)
  • 124 — Timeout (preset hung or took too long)
  • Other — Failure (check output.log)

Metrics in metrics.json:

  • iterations — How many event loop cycles
  • hats_activated — Which hats were triggered
  • events_published — Total events emitted
  • completed — Whether completion promise was reached

Hat Routing Performance

Critical: Validate that hats get fresh context per Tenet #1 ("Fresh Context Is Reliability").

What Good Looks Like

Each hat should execute in its own iteration:

Iter 1: Ralph → publishes starting event → STOPS
Iter 2: Hat A → does work → publishes next event → STOPS
Iter 3: Hat B → does work → publishes next event → STOPS
Iter 4: Hat C → does work → LOOP_COMPLETE
Red Flags (Same-Iteration Hat Switching)

BAD: Multiple hat personas in one iteration:

Iter 2: Ralph does Blue Team + Red Team + Fixer work
        ^^^ All in one bloated context!
Show full SKILL.md (265 more words)Show less
How to Check

1. Count iterations vs events in session.jsonl:

bash
# Count iterations
grep -c "_meta.loop_start\|ITERATION" .eval/logs/<preset>/latest/output.log

# Count events published
grep -c "bus.publish" .eval/logs/<preset>/latest/session.jsonl

Expected: iterations ≈ events published (one event per iteration) Bad sign: 2-3 iterations but 5+ events (all work in single iteration)

2. Check for same-iteration hat switching in output.log:

bash
grep -E "ITERATION|Now I need to perform|Let me put on|I'll switch to" \
    .eval/logs/<preset>/latest/output.log

Red flag: Hat-switching phrases WITHOUT an ITERATION separator between them.

3. Check event timestamps in session.jsonl:

bash
cat .eval/logs/<preset>/latest/session.jsonl | jq -r '.ts'

Red flag: Multiple events with identical timestamps (published in same iteration).

Routing Performance Triage
PatternDiagnosisAction
iterations ≈ events✅ GoodHat routing working
iterations << events⚠️ Same-iteration switchingCheck prompt has STOP instruction
iterations >> events⚠️ Recovery loopsAgent not publishing required events
0 events❌ BrokenEvents not being read from JSONL
Root Cause Checklist

If hat routing is broken:

  1. Check workflow prompt in hatless_ralph.rs:

    • Does it say "CRITICAL: STOP after publishing"?
    • Is the DELEGATE section clear about yielding control?
  2. Check hat instructions propagation:

    • Does HatInfo include instructions field?
    • Are instructions rendered in the ## HATS section?
  3. Check events context:

    • Is build_prompt(context) using the context parameter?
    • Does prompt include ## PENDING EVENTS section?

Autonomous Fix Workflow

After evaluation, delegate fixes to subagents:

Step 1: Triage Results

Read .eval/results/latest/SUMMARY.md and identify:

  • ❌ FAIL → Create code tasks for fixes
  • ⏱️ TIMEOUT → Investigate infinite loops
  • ⚠️ PARTIAL → Check for edge cases
Step 2: Dispatch Task Creation

For each issue, spawn a Task agent:

"Use /code-task-generator to create a task for fixing: [issue from evaluation]
Output to: .ralph/tasks/preset-fixes/"
Step 3: Dispatch Implementation

For each created task:

"Use /code-assist to implement: .ralph/tasks/preset-fixes/[task-file].code-task.md
Mode: auto"
Step 4: Re-evaluate
bash
./tools/evaluate-preset.sh <fixed-preset> claude

Prerequisites

  • yq (optional): For loading test tasks from YAML. Install: brew install yq
  • Cargo: Must be able to build Ralph
  • tools/evaluate-preset.sh — Single preset evaluation
  • tools/evaluate-all-presets.sh — Full suite evaluation
  • tools/preset-test-tasks.yml — Test task definitions
  • tools/preset-evaluation-findings.md — Manual findings doc
  • presets/ — The preset collection being evaluated

© mikeyobrien, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/evaluate-presets of mikeyobrien/ralph-orchestrator.

Open the folder on GitHubat commit edc2b32

Compare with similar skills

Evaluate Presets next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluate Presets compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluate Presets this skillmikeyobrien/ralph-orchestrator3.2k—~1.8kAutomated safety check: PassMIT
Arize Evaluatorgithub/awesome-copilot40k1 repos~8.1kAutomated safety check: NotesMIT
RalphYeachan-Heo/oh-my-claudecode40k—~7.5kAutomated safety check: PassMIT
Form Validationthedaviddias/Front-End-Checklist74k—~633Automated safety check: PassMIT
LLM Evaluationdavila7/claude-code-templates32k12 repos~3.5kAutomated safety check: PassMIT
Agent Evaluationsickn33/agentic-awesome-skills47k1 repos~2kAutomated safety check: PassMIT

Similar skills

  • Arize Evaluator

    github/awesome-copilot

    Official

    Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

    40k GitHub starsUsed in 1 repo~8.1k tokens
    AI & LLM EngineeringAuto-check: notes
  • Ralph

    Yeachan-Heo/oh-my-claudecode

    Self-referential loop until task completion with configurable verification reviewer

    40k GitHub stars~7.5k tokensUpdated today
    Product & Project ManagementAuto-check passed
  • Form Validation

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing templates, rendered HTML, or shared components related to Validate forms accessibly.

    74k GitHub stars~633 tokensUpdated 2 days ago
    Frontend & DesignAuto-check passed
  • LLM Evaluation

    davila7/claude-code-templates

    Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

    32k GitHub starsUsed in 12 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Evaluation

    sickn33/agentic-awesome-skills

    Evaluate agent behavior with versioned cases and explicit verifiers.

    47k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed
  • Validate

    agenticnotetaking/arscontexta

    Schema validation for notes. An agent skill from agenticnotetaking/arscontexta.

    3.5k GitHub stars~3k tokensUpdated 7 mo ago
    Frontend & DesignAuto-check passed

More from mikeyobrien/ralph-orchestrator

All 16 skills in this repo
  • Ralph Docs

    mikeyobrien/ralph-orchestrator

    Introspect, explain, and improve Ralph Orchestrator using its published llms.txt doc map.

    3.2k GitHub stars~1.5k tokensUpdated 4 days ago
    Auto-check passed
  • PR Demo

    mikeyobrien/ralph-orchestrator

    A skill your agent uses when creating animated demos (GIFs) for pull requests or documentation.

    3.2k GitHub stars~1.3k tokensUpdated 4 days ago
    Auto-check passed
  • Release Bump

    mikeyobrien/ralph-orchestrator

    A skill your agent uses when bumping ralph-orchestrator version for a new release, after fixes are committed and ready to publish

    3.2k GitHub stars~583 tokensUpdated 4 days ago
    Auto-check passed
  • Review PR

    mikeyobrien/ralph-orchestrator

    A skill your agent uses when asked to review a PR, run a code review loop, or invoke the ralph reviewer against a pull request number or GitHub URL

    3.2k GitHub stars~589 tokensUpdated 4 days ago
    Auto-check passed
  • Tui Debug In Pane

    mikeyobrien/ralph-orchestrator

    A skill your agent uses when you need to reproduce or debug TUI rendering issues (garbled output, broken streaming, layout corruption) by running ralph in a tmux split pane and capturing live output.

    3.2k GitHub stars~825 tokensUpdated 4 days ago
    Auto-check passed
  • Ralph Hats

    mikeyobrien/ralph-orchestrator

    Create, inspect, validate, explain, and improve Ralph hat collections.

    3.2k GitHub stars~691 tokensUpdated 4 days ago
    Auto-check passed

Questions about Evaluate Presets

What does Evaluate Presets do?

A skill your agent uses when testing Ralph's hat collection presets, validating preset configurations, or auditing the preset library for bugs and UX issues. Evaluate Presets is an agent skill from mikeyobrien/ralph-orchestrator. Use when testing Ralph's hat collection presets, validating preset configurations, or auditing the preset library for bugs and UX issues.

When should I use Evaluate Presets?

Evaluate Presets fits situations like: testing Ralphs hat collection presets; validating preset configurations; auditing the preset library for bugs and UX issues.

How do I install Evaluate Presets in Claude Code?

Run `npx skills add mikeyobrien/ralph-orchestrator --skill evaluate-presets -a claude-code`. Or copy the skill folder (.claude/skills/evaluate-presets in mikeyobrien/ralph-orchestrator) into .claude/skills/evaluate-presets in your project. Claude Code loads it when a task matches its description.

How do I install Evaluate Presets in Codex?

Run `npx skills add mikeyobrien/ralph-orchestrator --skill evaluate-presets -a codex`. Or copy the skill folder (.claude/skills/evaluate-presets in mikeyobrien/ralph-orchestrator) into .agents/skills/evaluate-presets in your project. Codex loads it when a task matches its description.

Can I use Evaluate Presets in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mikeyobrien/ralph-orchestrator --skill evaluate-presets -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluate-presets, .gemini/skills/evaluate-presets, .github/skills/evaluate-presets and .opencode/skills/evaluate-presets in your project.

What does Evaluate Presets need to run?

Going by SKILL.md and its folder, Evaluate Presets needs the command-line tools its instructions call (bash, jq and brew).

Does Evaluate Presets access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evaluate Presets safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evaluate Presets use?

Evaluate Presets is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluate Presets use?

About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evaluate Presets?

Skills that share tags, products or a category with Evaluate Presets: Arize Evaluator (github/awesome-copilot, 40k stars), Ralph (Yeachan-Heo/oh-my-claudecode, 40k stars), Form Validation (thedaviddias/Front-End-Checklist, 74k stars) and LLM Evaluation (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluate Presets?

mikeyobrien (a GitHub user) maintains it in mikeyobrien/ralph-orchestrator, which has 3,169 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on October 5, 2026.

Source: mikeyobrien/ralph-orchestrator on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.