Arize Evaluator
github/awesome-copilot
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…
A skill your agent uses when testing Ralph's hat collection presets, validating preset configurations, or auditing the preset library for bugs and UX issues.
$ npx skills add mikeyobrien/ralph-orchestrator --skill evaluate-presets -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mikeyobrien/ralph-orchestrator evaluate-presets --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mikeyobrien/ralph-orchestrator.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/evaluate-presets .claude/skills/evaluate-presets && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "evaluate-presets" agent skill from https://github.com/mikeyobrien/ralph-orchestrator/tree/main/.claude/skills/evaluate-presets into .claude/skills/evaluate-presets/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-presets", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mikeyobrien/ralph-orchestrator/tree/main/.claude/skills/evaluate-presetsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mikeyobrien/ralph-orchestrator --skill evaluate-presets -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mikeyobrien/ralph-orchestrator evaluate-presets --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mikeyobrien/ralph-orchestrator.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/evaluate-presets .agents/skills/evaluate-presets && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "evaluate-presets" agent skill from https://github.com/mikeyobrien/ralph-orchestrator/tree/main/.claude/skills/evaluate-presets into .agents/skills/evaluate-presets/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-presets", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mikeyobrien/ralph-orchestrator --skill evaluate-presets -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mikeyobrien/ralph-orchestrator evaluate-presets --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mikeyobrien/ralph-orchestrator.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/evaluate-presets .cursor/skills/evaluate-presets && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "evaluate-presets" agent skill from https://github.com/mikeyobrien/ralph-orchestrator/tree/main/.claude/skills/evaluate-presets into .cursor/skills/evaluate-presets/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-presets", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mikeyobrien/ralph-orchestrator.git --path .claude/skills/evaluate-presets--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mikeyobrien/ralph-orchestrator --skill evaluate-presets -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mikeyobrien/ralph-orchestrator evaluate-presets --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mikeyobrien/ralph-orchestrator.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/evaluate-presets .gemini/skills/evaluate-presets && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "evaluate-presets" agent skill from https://github.com/mikeyobrien/ralph-orchestrator/tree/main/.claude/skills/evaluate-presets into .gemini/skills/evaluate-presets/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-presets", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mikeyobrien/ralph-orchestrator evaluate-presetsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mikeyobrien/ralph-orchestrator --skill evaluate-presets -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mikeyobrien/ralph-orchestrator.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/evaluate-presets .github/skills/evaluate-presets && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "evaluate-presets" agent skill from https://github.com/mikeyobrien/ralph-orchestrator/tree/main/.claude/skills/evaluate-presets into .github/skills/evaluate-presets/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-presets", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mikeyobrien/ralph-orchestrator --skill evaluate-presets -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mikeyobrien/ralph-orchestrator evaluate-presets --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mikeyobrien/ralph-orchestrator.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/evaluate-presets .opencode/skills/evaluate-presets && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "evaluate-presets" agent skill from https://github.com/mikeyobrien/ralph-orchestrator/tree/main/.claude/skills/evaluate-presets into .opencode/skills/evaluate-presets/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-presets", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
evaluate-presetsA skill your agent uses when testing Ralph's hat collection presets, validating preset configurations, or auditing the preset library for bugs and UX issues.
Evaluate Presets is an agent skill from mikeyobrien/ralph-orchestrator. Use when testing Ralph's hat collection presets, validating preset configurations, or auditing the preset library for bugs and UX issues.
Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: An improved implementation of the Ralph Wiggum technique for autonomous AI agent orchestration. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit edc2b32. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
bashjqbrewFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Evaluate Presets loads about 1.8k tokens when it runs. Until then it costs about 39 tokens; SKILL.md has 625 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mikeyobrien/ralph-orchestrator at commit edc2b32, republished under its MIT licence (© mikeyobrien). 625 words, ~1,828 tokens.
.claude/skills/evaluate-presets/SKILL.md (or your agent's skills folder).Systematically test all hat collection presets using shell scripts. Direct CLI invocation—no meta-orchestration complexity.
Evaluate a single preset:
./tools/evaluate-preset.sh tdd-red-green claudeEvaluate all presets:
./tools/evaluate-all-presets.sh claudeArguments:
.yml extension)claude or kiro, defaults to claude)IMPORTANT: When invoking these scripts via the Bash tool, use these settings:
timeout: 600000 (10 minutes max) and run_in_background: truetimeout: 600000 (10 minutes max) and run_in_background: trueSince preset evaluations can run for hours (especially the full suite), always run in background mode and use the TaskOutput tool to check progress periodically.
Example invocation pattern:
Bash tool with:
command: "./tools/evaluate-preset.sh tdd-red-green claude"
timeout: 600000
run_in_background: trueAfter launching, use TaskOutput with block: false to check status without waiting for completion.
evaluate-preset.shtools/preset-test-tasks.yml (if yq available)--record-session for metrics captureOutput structure:
.eval/
├── logs/<preset>/<timestamp>/
│ ├── output.log # Full stdout/stderr
│ ├── session.jsonl # Recorded session
│ ├── metrics.json # Extracted metrics
│ ├── environment.json # Runtime environment
│ └── merged-config.yml # Config used
└── logs/<preset>/latest -> <timestamp>evaluate-all-presets.shRuns all 12 presets sequentially and generates a summary:
.eval/results/<suite-id>/
├── SUMMARY.md # Markdown report
├── <preset>.json # Per-preset metrics
└── latest -> <suite-id>| Preset | Test Task |
|---|---|
tdd-red-green | Add is_palindrome() function |
adversarial-review | Review user input handler for security |
socratic-learning | Understand HatRegistry |
spec-driven | Specify and implement StringUtils::truncate() |
mob-programming | Implement a Stack data structure |
scientific-method | Debug failing mock test assertion |
code-archaeology | Understand history of config.rs |
performance-optimization | Profile hat matching |
api-design | Design a Cache trait |
documentation-first | Document RateLimiter |
incident-response | Respond to "tests failing in CI" |
migration-safety | Plan v1 to v2 config migration |
Exit codes from evaluate-preset.sh:
0 — Success (LOOP_COMPLETE reached)124 — Timeout (preset hung or took too long)output.log)Metrics in metrics.json:
iterations — How many event loop cycleshats_activated — Which hats were triggeredevents_published — Total events emittedcompleted — Whether completion promise was reachedCritical: Validate that hats get fresh context per Tenet #1 ("Fresh Context Is Reliability").
Each hat should execute in its own iteration:
Iter 1: Ralph → publishes starting event → STOPS
Iter 2: Hat A → does work → publishes next event → STOPS
Iter 3: Hat B → does work → publishes next event → STOPS
Iter 4: Hat C → does work → LOOP_COMPLETEBAD: Multiple hat personas in one iteration:
Iter 2: Ralph does Blue Team + Red Team + Fixer work
^^^ All in one bloated context!1. Count iterations vs events in session.jsonl:
# Count iterations
grep -c "_meta.loop_start\|ITERATION" .eval/logs/<preset>/latest/output.log
# Count events published
grep -c "bus.publish" .eval/logs/<preset>/latest/session.jsonlExpected: iterations ≈ events published (one event per iteration) Bad sign: 2-3 iterations but 5+ events (all work in single iteration)
2. Check for same-iteration hat switching in output.log:
grep -E "ITERATION|Now I need to perform|Let me put on|I'll switch to" \
.eval/logs/<preset>/latest/output.logRed flag: Hat-switching phrases WITHOUT an ITERATION separator between them.
3. Check event timestamps in session.jsonl:
cat .eval/logs/<preset>/latest/session.jsonl | jq -r '.ts'Red flag: Multiple events with identical timestamps (published in same iteration).
| Pattern | Diagnosis | Action |
|---|---|---|
| iterations ≈ events | ✅ Good | Hat routing working |
| iterations << events | ⚠️ Same-iteration switching | Check prompt has STOP instruction |
| iterations >> events | ⚠️ Recovery loops | Agent not publishing required events |
| 0 events | ❌ Broken | Events not being read from JSONL |
If hat routing is broken:
Check workflow prompt in hatless_ralph.rs:
Check hat instructions propagation:
HatInfo include instructions field?## HATS section?Check events context:
build_prompt(context) using the context parameter?## PENDING EVENTS section?After evaluation, delegate fixes to subagents:
Read .eval/results/latest/SUMMARY.md and identify:
❌ FAIL → Create code tasks for fixes⏱️ TIMEOUT → Investigate infinite loops⚠️ PARTIAL → Check for edge casesFor each issue, spawn a Task agent:
"Use /code-task-generator to create a task for fixing: [issue from evaluation]
Output to: .ralph/tasks/preset-fixes/"For each created task:
"Use /code-assist to implement: .ralph/tasks/preset-fixes/[task-file].code-task.md
Mode: auto"./tools/evaluate-preset.sh <fixed-preset> claudebrew install yqtools/evaluate-preset.sh — Single preset evaluationtools/evaluate-all-presets.sh — Full suite evaluationtools/preset-test-tasks.yml — Test task definitionstools/preset-evaluation-findings.md — Manual findings docpresets/ — The preset collection being evaluated© mikeyobrien, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/evaluate-presets of mikeyobrien/ralph-orchestrator.
Open the folder on GitHubat commit edc2b32
Evaluate Presets next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Evaluate Presets this skillmikeyobrien/ralph-orchestrator | 3.2k | — | ~1.8k | Automated safety check: Pass | MIT | |
| Arize Evaluatorgithub/awesome-copilot | 40k | 1 repos | ~8.1k | Automated safety check: Notes | MIT | |
| RalphYeachan-Heo/oh-my-claudecode | 40k | — | ~7.5k | Automated safety check: Pass | MIT | |
| Form Validationthedaviddias/Front-End-Checklist | 74k | — | ~633 | Automated safety check: Pass | MIT | |
| LLM Evaluationdavila7/claude-code-templates | 32k | 12 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Agent Evaluationsickn33/agentic-awesome-skills | 47k | 1 repos | ~2k | Automated safety check: Pass | MIT |
github/awesome-copilot
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…
Yeachan-Heo/oh-my-claudecode
Self-referential loop until task completion with configurable verification reviewer
thedaviddias/Front-End-Checklist
A skill your agent uses when reviewing templates, rendered HTML, or shared components related to Validate forms accessibly.
davila7/claude-code-templates
Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.
sickn33/agentic-awesome-skills
Evaluate agent behavior with versioned cases and explicit verifiers.
agenticnotetaking/arscontexta
Schema validation for notes. An agent skill from agenticnotetaking/arscontexta.
mikeyobrien/ralph-orchestrator
Introspect, explain, and improve Ralph Orchestrator using its published llms.txt doc map.
mikeyobrien/ralph-orchestrator
A skill your agent uses when creating animated demos (GIFs) for pull requests or documentation.
mikeyobrien/ralph-orchestrator
A skill your agent uses when bumping ralph-orchestrator version for a new release, after fixes are committed and ready to publish
mikeyobrien/ralph-orchestrator
A skill your agent uses when asked to review a PR, run a code review loop, or invoke the ralph reviewer against a pull request number or GitHub URL
mikeyobrien/ralph-orchestrator
A skill your agent uses when you need to reproduce or debug TUI rendering issues (garbled output, broken streaming, layout corruption) by running ralph in a tmux split pane and capturing live output.
mikeyobrien/ralph-orchestrator
Create, inspect, validate, explain, and improve Ralph hat collections.
A skill your agent uses when testing Ralph's hat collection presets, validating preset configurations, or auditing the preset library for bugs and UX issues. Evaluate Presets is an agent skill from mikeyobrien/ralph-orchestrator. Use when testing Ralph's hat collection presets, validating preset configurations, or auditing the preset library for bugs and UX issues.
Evaluate Presets fits situations like: testing Ralphs hat collection presets; validating preset configurations; auditing the preset library for bugs and UX issues.
Run `npx skills add mikeyobrien/ralph-orchestrator --skill evaluate-presets -a claude-code`. Or copy the skill folder (.claude/skills/evaluate-presets in mikeyobrien/ralph-orchestrator) into .claude/skills/evaluate-presets in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mikeyobrien/ralph-orchestrator --skill evaluate-presets -a codex`. Or copy the skill folder (.claude/skills/evaluate-presets in mikeyobrien/ralph-orchestrator) into .agents/skills/evaluate-presets in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mikeyobrien/ralph-orchestrator --skill evaluate-presets -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluate-presets, .gemini/skills/evaluate-presets, .github/skills/evaluate-presets and .opencode/skills/evaluate-presets in your project.
Going by SKILL.md and its folder, Evaluate Presets needs the command-line tools its instructions call (bash, jq and brew).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Evaluate Presets is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Evaluate Presets: Arize Evaluator (github/awesome-copilot, 40k stars), Ralph (Yeachan-Heo/oh-my-claudecode, 40k stars), Form Validation (thedaviddias/Front-End-Checklist, 74k stars) and LLM Evaluation (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mikeyobrien (a GitHub user) maintains it in mikeyobrien/ralph-orchestrator, which has 3,169 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on October 5, 2026.
Source: mikeyobrien/ralph-orchestrator on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.