Tui Validate
mikeyobrien/ralph-orchestrator
Validates Terminal User Interface (TUI) output using freeze for screenshot capture and LLM-as-judge for semantic validation.
Measure whether a change makes Chat's styling replies better, with yarn eval:chat — it runs styling requests on recorded real pages for two versions of Stylebot and measures on the page whether each…
$ npx skills add ankit/stylebot --skill eval-chat -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ankit/stylebot eval-chat --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ankit/stylebot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/eval-chat .claude/skills/eval-chat && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-chat" agent skill from https://github.com/ankit/stylebot/tree/v4/.agents/skills/eval-chat into .claude/skills/eval-chat/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-chat", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ankit/stylebot/tree/v4/.agents/skills/eval-chatType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ankit/stylebot --skill eval-chat -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ankit/stylebot eval-chat --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ankit/stylebot.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/eval-chat .agents/skills/eval-chat && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-chat" agent skill from https://github.com/ankit/stylebot/tree/v4/.agents/skills/eval-chat into .agents/skills/eval-chat/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-chat", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ankit/stylebot --skill eval-chat -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ankit/stylebot eval-chat --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ankit/stylebot.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/eval-chat .cursor/skills/eval-chat && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-chat" agent skill from https://github.com/ankit/stylebot/tree/v4/.agents/skills/eval-chat into .cursor/skills/eval-chat/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-chat", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ankit/stylebot.git --path .agents/skills/eval-chat--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ankit/stylebot --skill eval-chat -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ankit/stylebot eval-chat --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ankit/stylebot.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/eval-chat .gemini/skills/eval-chat && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-chat" agent skill from https://github.com/ankit/stylebot/tree/v4/.agents/skills/eval-chat into .gemini/skills/eval-chat/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-chat", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ankit/stylebot eval-chatInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ankit/stylebot --skill eval-chat -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ankit/stylebot.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/eval-chat .github/skills/eval-chat && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-chat" agent skill from https://github.com/ankit/stylebot/tree/v4/.agents/skills/eval-chat into .github/skills/eval-chat/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-chat", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ankit/stylebot --skill eval-chat -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ankit/stylebot eval-chat --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ankit/stylebot.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/eval-chat .opencode/skills/eval-chat && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-chat" agent skill from https://github.com/ankit/stylebot/tree/v4/.agents/skills/eval-chat into .opencode/skills/eval-chat/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-chat", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-chatMeasure whether a change makes Chat's styling replies better, with yarn eval:chat — it runs styling requests on recorded real pages for two versions of Stylebot and measures on the page whether each…
Eval Chat is an agent skill from ankit/stylebot. Measure whether a change makes Chat's styling replies better, with yarn eval:chat — it runs styling requests on recorded real pages for two versions of Stylebot and measures on the page whether each request was met. Use after changing Chat's prompt, the page outline or CSS context it sends, or how replies are applied or checked, before claiming the change helps; when the user asks "is this better?", "run the evals", or wants numbers for a PR.
Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Frontend & Design, covering LLM evaluation. The repository describes itself as: Change the appearance of the web instantly. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 96a61f5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
yarnFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use yarn, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval Chat loads about 1.1k tokens when it runs. Until then it costs about 115 tokens; SKILL.md has 629 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ankit/stylebot at commit 96a61f5, republished under its MIT licence (© ankit). 629 words, ~1,096 tokens.
.claude/skills/eval-chat/SKILL.md (or your agent's skills folder).Unit tests show Chat's plumbing works, not that its replies look better. yarn eval:chat
measures that: each case is a request on a recorded page ("Make this page an everforest
theme with Fira Code as typography" on Hacker News), run for a base and a head version,
then measured on the page (a column's width, a sidebar hidden, a font, a color) and shown
in screenshots. Model calls go through headless Claude Code on the user's subscription,
not an API key. docs/chat-evals.md describes the harness; this is how to
use it well.
From the checkout with the change, after yarn install:
yarn eval:chat --base HEAD --head .--base HEAD --head .) while iterating, so the
numbers isolate this change; compare against v4 for a PR's description.--runs 3. To compare models or thinking, give head its own setup with --head-model
or --head-thinking.Results in …; it ends by saying how many model
calls it made and how long it took. Results are cached by the code that produced them,
so an unchanged base costs nothing the second time.Open summary.md in the results folder.
docs/chat-evals.md).yarn eval:chat --references --cases <id> first: no
model calls, and every check should pass with the case's reference stylesheet. Write a
reference for a case that doesn't have one yet.What moved the scores, roughly in order:
bgcolor, a design system's base palette buried under
thousands of CSS variables.In the PR description, give the comparison against v4: the overall row (checks,
unreadable, zero-match, cost), the model setup and runs, and any case that got worse.
© ankit, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/eval-chat of ankit/stylebot.
Open the folder on GitHubat commit 96a61f5
Eval Chat next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval Chat this skillankit/stylebot | 1.6k | — | ~1.1k | Automated safety check: Pass | MIT | |
| Tui Validatemikeyobrien/ralph-orchestrator | 3.2k | — | ~3k | Automated safety check: Pass | MIT | |
| LLM Trace Review Interfaceai-evals-course/evals-skills | 1.5k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| EvaluationPrimeIntellect-ai/prime-envs | 130 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Eval Guidemicrosoft/eval-guide | 138 | — | ~22k | Automated safety check: Warn | MIT | |
| Write A Specdifferent-ai/openwork | 24k | — | ~3.3k | Automated safety check: Pass | Custom licence |
mikeyobrien/ralph-orchestrator
Validates Terminal User Interface (TUI) output using freeze for screenshot capture and LLM-as-judge for semantic validation.
ai-evals-course/evals-skills
Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data.
PrimeIntellect-ai/prime-envs
Install and run a verifiers environment — smoke testing during development and full benchmark evals.
microsoft/eval-guide
Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately.
different-ai/openwork
Write or extend an E2E journey spec in evals/specs that proves a PR's change to a human reviewer.
JuliusBrussee/caveman
Reads the state and results of Caveman Cloud experiments and reports one recommendation or a block, without changing an experiment's lifecycle itself.
ankit/stylebot
Cut a Stylebot release — CHANGELOG entry, version bump, release PR, tag, and store zips.
ankit/stylebot
Translate Stylebot's locale strings — fill in strings that English has and other locales lack, retranslate English strings that changed, or add a whole new locale (e.g.
ankit/stylebot
Launch Stylebot in a real, headed Chrome window (via yarn dev:chrome) so the user can manually eyeball a change — in whatever checkout the current session is in, worktree or main.
ankit/stylebot
Restyle websites in the user's browser with the Stylebot extension's stylebot CLI — dark modes, themes, fonts, bigger text, hiding clutter, fixing a layout, or editing and managing the CSS and…
Categories
Measure whether a change makes Chat's styling replies better, with yarn eval:chat — it runs styling requests on recorded real pages for two versions of Stylebot and measures on the page whether each…. Eval Chat is an agent skill from ankit/stylebot. Measure whether a change makes Chat's styling replies better, with yarn eval:chat — it runs styling requests on recorded real pages for two versions of Stylebot and measures on the page whether each request was met.
Eval Chat fits situations like: asks is this better?; wants numbers for a PR.
Run `npx skills add ankit/stylebot --skill eval-chat -a claude-code`. Or copy the skill folder (.agents/skills/eval-chat in ankit/stylebot) into .claude/skills/eval-chat in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ankit/stylebot --skill eval-chat -a codex`. Or copy the skill folder (.agents/skills/eval-chat in ankit/stylebot) into .agents/skills/eval-chat in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ankit/stylebot --skill eval-chat -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-chat, .gemini/skills/eval-chat, .github/skills/eval-chat and .opencode/skills/eval-chat in your project.
Going by SKILL.md and its folder, Eval Chat needs the command-line tools its instructions call (yarn).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Eval Chat is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.1k tokens (SKILL.md is roughly 4.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Eval Chat: Tui Validate (mikeyobrien/ralph-orchestrator, 3.2k stars), LLM Trace Review Interface (ai-evals-course/evals-skills, 1.5k stars), Evaluation (PrimeIntellect-ai/prime-envs, 130 stars) and Eval Guide (microsoft/eval-guide, 138 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ankit (a GitHub user) maintains it in ankit/stylebot, which has 1,647 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 8, 2026.
Source: ankit/stylebot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.