Eval Agentic Launch
open-thoughts/OpenThoughts-Agent
Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy…
Validates Terminal User Interface (TUI) output using freeze for screenshot capture and LLM-as-judge for semantic validation.
$ npx skills add mikeyobrien/ralph-orchestrator --skill tui-validate -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mikeyobrien/ralph-orchestrator tui-validate --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mikeyobrien/ralph-orchestrator.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/tui-validate .claude/skills/tui-validate && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tui-validate" agent skill from https://github.com/mikeyobrien/ralph-orchestrator/tree/main/.claude/skills/tui-validate into .claude/skills/tui-validate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tui-validate", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mikeyobrien/ralph-orchestrator/tree/main/.claude/skills/tui-validateType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mikeyobrien/ralph-orchestrator --skill tui-validate -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mikeyobrien/ralph-orchestrator tui-validate --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mikeyobrien/ralph-orchestrator.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/tui-validate .agents/skills/tui-validate && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tui-validate" agent skill from https://github.com/mikeyobrien/ralph-orchestrator/tree/main/.claude/skills/tui-validate into .agents/skills/tui-validate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tui-validate", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mikeyobrien/ralph-orchestrator --skill tui-validate -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mikeyobrien/ralph-orchestrator tui-validate --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mikeyobrien/ralph-orchestrator.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/tui-validate .cursor/skills/tui-validate && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tui-validate" agent skill from https://github.com/mikeyobrien/ralph-orchestrator/tree/main/.claude/skills/tui-validate into .cursor/skills/tui-validate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tui-validate", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mikeyobrien/ralph-orchestrator.git --path .claude/skills/tui-validate--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mikeyobrien/ralph-orchestrator --skill tui-validate -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mikeyobrien/ralph-orchestrator tui-validate --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mikeyobrien/ralph-orchestrator.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/tui-validate .gemini/skills/tui-validate && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tui-validate" agent skill from https://github.com/mikeyobrien/ralph-orchestrator/tree/main/.claude/skills/tui-validate into .gemini/skills/tui-validate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tui-validate", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mikeyobrien/ralph-orchestrator tui-validateInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mikeyobrien/ralph-orchestrator --skill tui-validate -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mikeyobrien/ralph-orchestrator.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/tui-validate .github/skills/tui-validate && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tui-validate" agent skill from https://github.com/mikeyobrien/ralph-orchestrator/tree/main/.claude/skills/tui-validate into .github/skills/tui-validate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tui-validate", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mikeyobrien/ralph-orchestrator --skill tui-validate -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mikeyobrien/ralph-orchestrator tui-validate --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mikeyobrien/ralph-orchestrator.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/tui-validate .opencode/skills/tui-validate && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tui-validate" agent skill from https://github.com/mikeyobrien/ralph-orchestrator/tree/main/.claude/skills/tui-validate into .opencode/skills/tui-validate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tui-validate", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tui-validateValidates Terminal User Interface (TUI) output using freeze for screenshot capture and LLM-as-judge for semantic validation.
Tui Validate is an agent skill from mikeyobrien/ralph-orchestrator. Validates Terminal User Interface (TUI) output using freeze for screenshot capture and LLM-as-judge for semantic validation. Supports both visual (PNG/SVG) and text-based validation modes.
Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM evaluation. It works with tmux. The repository describes itself as: An improved implementation of the Ralph Wiggum technique for autonomous AI agent orchestration. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit edc2b32. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
brewgoFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Tui Validate loads about 3k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 903 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mikeyobrien/ralph-orchestrator at commit edc2b32, republished under its MIT licence (© mikeyobrien). 903 words, ~3,024 tokens.
.claude/skills/tui-validate/SKILL.md (or your agent's skills folder).This skill validates Terminal User Interface (TUI) applications by capturing their output and using LLM-as-judge for semantic validation. It leverages freeze from Charmbracelet for high-fidelity terminal screenshots and provides structured validation criteria.
Philosophy: Rather than brittle string matching, this skill uses semantic understanding to validate that TUI output "looks right" - checking layout, content presence, and visual hierarchy without breaking on minor formatting changes.
Required:
freeze CLI tool installed (brew install charmbracelet/tap/freeze)tmux for interactive TUI capture (optional, for live applications)Verification:
# Check freeze is installed
freeze --version
# Check tmux is installed (for interactive capture)
tmux -Vtarget (required): What to validate. One of:
file:<path> - ANSI output file to validatecommand:<cmd> - Command to execute and capturetmux:<session> - Live tmux session to capturebuffer:<text> - Raw text/ANSI to validatecriteria (required): Validation criteria. Can be:
output_format (optional, default: "svg"): Screenshot format
svg - Vector format, best for documentationpng - Raster format, best for visual difftext - Text-only extraction, fastestsave_screenshot (optional, default: false): Whether to save the screenshot
{target_name}.{format} in current directoryjudge_mode (optional, default: "semantic"): Validation approach
semantic - LLM judges based on meaning and layoutstrict - Also checks exact content presencevisual - Requires PNG, checks visual appearanceralph-headerValidates Ralph TUI header component:
[iter N] or [iter N/M] formatMM:SS format▶ auto or ⏸ paused)[SCROLL]idle: Nsralph-footerValidates Ralph TUI footer component:
◉ active, ◯ idle, or ■ done)ralph-fullValidates complete Ralph TUI layout:
tui-basicGeneric TUI validation:
Capture TUI output based on target type:
For file targets:
freeze {file_path} -o /tmp/tui-capture.{format}For command targets:
freeze --execute "{command}" -o /tmp/tui-capture.{format}For tmux targets:
tmux capture-pane -pet {session} | freeze -o /tmp/tui-capture.{format}For buffer targets:
echo "{buffer}" | freeze -o /tmp/tui-capture.{format}Constraints:
--theme base16 for consistent rendering--width and --heightExtract content for LLM analysis:
For text/semantic validation:
text, use the captured text directlysvg or png, also capture text version for content analysisFor visual validation:
Constraints:
visualApply LLM-as-judge with the appropriate criteria:
Semantic Validation Prompt Template:
Analyze this terminal UI output and determine if it meets the following criteria:
CRITERIA:
{criteria_description}
TERMINAL OUTPUT:
{captured_text}
Evaluate each criterion and provide:
1. PASS or FAIL for each requirement
2. Brief explanation for any failures
3. Overall verdict: PASS or FAIL
Be lenient on exact formatting but strict on:
- Required content presence
- Logical layout and hierarchy
- No rendering errors or artifactsVisual Validation Prompt Template (with image):
Examine this terminal screenshot and validate:
CRITERIA:
{criteria_description}
Check for:
1. Visual hierarchy and layout
2. Color coding correctness
3. No rendering artifacts or broken characters
4. Proper alignment and spacing
Verdict: PASS or FAIL with explanationConstraints:
Report validation results:
On PASS:
✅ TUI Validation PASSED
Criteria: {criteria_name}
Target: {target}
Mode: {judge_mode}
All requirements satisfied.
{optional_notes}On FAIL:
❌ TUI Validation FAILED
Criteria: {criteria_name}
Target: {target}
Mode: {judge_mode}
Issues found:
- {issue_1}
- {issue_2}
Screenshot saved: {path_if_saved}Constraints:
Input:
/tui-validate file:test_output.txt criteria:ralph-headerProcess:
test_output.txt containing ANSI outputfreeze test_output.txt -o /tmp/capture.svgralph-header criteria via LLM judgeInput:
/tui-validate tmux:ralph-session criteria:ralph-full save_screenshot:trueProcess:
tmux capture-pane -pet ralph-session | freeze -o ralph-session.svgtmux capture-pane -pet ralph-session > /tmp/text.txtralph-full criteria checking header, content, and footerralph-session.svgInput:
/tui-validate command:"cargo run --example tui_demo" criteria:"Shows a bordered box with 'Hello World' text centered inside" output_format:png judge_mode:visualProcess:
freeze --execute "cargo run --example tui_demo" -o /tmp/capture.pngInput:
/tui-validate buffer:"[iter 3/10] 04:32 | 🔨 Builder | ▶ auto" criteria:ralph-header output_format:textProcess:
name: ralph-header
description: Ralph TUI header component validation
requirements:
- name: iteration_counter
description: Shows iteration in [iter N] or [iter N/M] format
required: true
pattern: '\[iter \d+(/\d+)?\]'
- name: elapsed_time
description: Shows elapsed time in MM:SS format
required: true
pattern: '\d{2}:\d{2}'
- name: hat_indicator
description: Shows current hat with emoji prefix
required: true
examples: ["🔨 Builder", "📋 Planner", "🎯 Executor"]
- name: mode_indicator
description: Shows loop mode status
required: true
values: ["▶ auto", "⏸ paused"]
- name: scroll_indicator
description: Shows [SCROLL] when in scroll mode
required: false
pattern: '\[SCROLL\]'
- name: idle_countdown
description: Shows idle timeout when present
required: false
pattern: 'idle: \d+s'name: ralph-footer
description: Ralph TUI footer component validation
requirements:
- name: activity_indicator
description: Shows current activity state
required: true
values: ["◉ active", "◯ idle", "■ done"]
- name: event_topic
description: Shows last event topic
required: false
examples: ["task.start", "build.done", "loop.terminate"]
- name: search_display
description: Shows search query and match count when searching
required: false
pattern: 'Search: .+ \d+/\d+'name: ralph-full
description: Complete Ralph TUI layout validation
requirements:
- name: header_section
description: Header at top with iteration, time, hat, and mode
required: true
references: ralph-header
- name: content_section
description: Main terminal content area
required: true
checks:
- Has visible content or is ready for content
- Properly bounded between header and footer
- name: footer_section
description: Footer at bottom with activity status
required: true
references: ralph-footer
- name: visual_hierarchy
description: Clear visual separation between sections
required: true
checks:
- Borders or spacing between sections
- Consistent width across sections# macOS
brew install charmbracelet/tap/freeze
# Linux (via Go)
go install github.com/charmbracelet/freeze@latest
# Verify installation
freeze --versiontmux list-sessionstmux list-panes -t {session}tmux capture-pane -pet {session}:{pane}--theme flag for consistent colorsstrict mode for exact matching requirementssemantic mode for layout/presence checkingThis skill can be integrated into test suites:
// In tests/tui_validation.rs
#[test]
#[ignore] // Run with: cargo test -- --ignored
fn validate_header_rendering() {
// 1. Render header to buffer
let output = render_header_to_string(&test_state);
// 2. Save to temp file
std::fs::write("/tmp/header_test.txt", &output).unwrap();
// 3. Run tui-validate skill (via CLI or programmatic)
// /tui-validate file:/tmp/header_test.txt criteria:ralph-header
// 4. Assert validation passed
}© mikeyobrien, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/tui-validate of mikeyobrien/ralph-orchestrator.
Open the folder on GitHubat commit edc2b32
Tui Validate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Tui Validate this skillmikeyobrien/ralph-orchestrator | 3.2k | — | ~3k | Automated safety check: Pass | MIT | |
| Eval Agentic Launchopen-thoughts/OpenThoughts-Agent | 301 | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | |
| LLM Trace Review Interfaceai-evals-course/evals-skills | 1.5k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| 1passwordtrpc-group/trpc-agent-go | 1.8k | 15 repos | ~656 | Automated safety check: Pass | Apache-2.0 | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Azure AI Projects Python SDKmicrosoft/skills | 3.1k | 6 repos | ~2.8k | Automated safety check: Pass | MIT |
open-thoughts/OpenThoughts-Agent
Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy…
ai-evals-course/evals-skills
Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data.
trpc-group/trpc-agent-go
Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
microsoft/skills
Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.
Jeffallan/claude-skills
Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.
mikeyobrien/ralph-orchestrator
Introspect, explain, and improve Ralph Orchestrator using its published llms.txt doc map.
mikeyobrien/ralph-orchestrator
A skill your agent uses when testing Ralph's hat collection presets, validating preset configurations, or auditing the preset library for bugs and UX issues.
mikeyobrien/ralph-orchestrator
A skill your agent uses when creating animated demos (GIFs) for pull requests or documentation.
mikeyobrien/ralph-orchestrator
A skill your agent uses when bumping ralph-orchestrator version for a new release, after fixes are committed and ready to publish
mikeyobrien/ralph-orchestrator
A skill your agent uses when asked to review a PR, run a code review loop, or invoke the ralph reviewer against a pull request number or GitHub URL
mikeyobrien/ralph-orchestrator
A skill your agent uses when you need to reproduce or debug TUI rendering issues (garbled output, broken streaming, layout corruption) by running ralph in a tmux split pane and capturing live output.
Works with
Categories
Validates Terminal User Interface (TUI) output using freeze for screenshot capture and LLM-as-judge for semantic validation. Tui Validate is an agent skill from mikeyobrien/ralph-orchestrator. Validates Terminal User Interface (TUI) output using freeze for screenshot capture and LLM-as-judge for semantic validation.
Tui Validate fits situations like: tasks that involve LLM evaluation.
Run `npx skills add mikeyobrien/ralph-orchestrator --skill tui-validate -a claude-code`. Or copy the skill folder (.claude/skills/tui-validate in mikeyobrien/ralph-orchestrator) into .claude/skills/tui-validate in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mikeyobrien/ralph-orchestrator --skill tui-validate -a codex`. Or copy the skill folder (.claude/skills/tui-validate in mikeyobrien/ralph-orchestrator) into .agents/skills/tui-validate in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mikeyobrien/ralph-orchestrator --skill tui-validate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tui-validate, .gemini/skills/tui-validate, .github/skills/tui-validate and .opencode/skills/tui-validate in your project.
Going by SKILL.md and its folder, Tui Validate needs the command-line tools its instructions call (brew and go).
SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Tui Validate is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Tui Validate: Eval Agentic Launch (open-thoughts/OpenThoughts-Agent, 301 stars), LLM Trace Review Interface (ai-evals-course/evals-skills, 1.5k stars), 1password (trpc-group/trpc-agent-go, 1.8k stars) and LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mikeyobrien (a GitHub user) maintains it in mikeyobrien/ralph-orchestrator, which has 3,167 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on October 5, 2026.
Source: mikeyobrien/ralph-orchestrator on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.