Agent skill

Eval Chat

by ankit in ankit/stylebot

Measure whether a change makes Chat's styling replies better, with yarn eval:chat — it runs styling requests on recorded real pages for two versions of Stylebot and measures on the page whether each…

MITAuto-check passedFrontend & Design

Install Eval Chat

skills CLI
$ npx skills add ankit/stylebot --skill eval-chat -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ankit/stylebot eval-chat --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ankit/stylebot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/eval-chat .claude/skills/eval-chat && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-chat
GitHub stars
1.6k
Token cost
~1.1k tokens
SKILL.md length
629 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Measure whether a change makes Chat's styling replies better, with yarn eval:chat — it runs styling requests on recorded real pages for two versions of Stylebot and measures on the page whether each…

  • Works in 4 steps: Run it → Read it → Improve what it finds → …
  • Asks is this better?
  • SKILL.md covers 1. Run it, 2. Read it, 3. Improve what it finds and 4. Report it
  • Calls yarn

What it does

Eval Chat is an agent skill from ankit/stylebot. Measure whether a change makes Chat's styling replies better, with yarn eval:chat — it runs styling requests on recorded real pages for two versions of Stylebot and measures on the page whether each request was met. Use after changing Chat's prompt, the page outline or CSS context it sends, or how replies are applied or checked, before claiming the change helps; when the user asks "is this better?", "run the evals", or wants numbers for a PR.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Frontend & Design, covering LLM evaluation. The repository describes itself as: Change the appearance of the web instantly. The licence is MIT.

When your agent uses it

  • Asks is this better?
  • Wants numbers for a PR

Example prompts

  • “is this better?”
  • “run the evals”
  • “/eval-chat”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Run it
  2. Read it
  3. Improve what it finds
  4. Report it

What it can do on your machine

Read from SKILL.md and the folder at commit 96a61f5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • yarn

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use yarn, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Chat loads about 1.1k tokens when it runs. Until then it costs about 115 tokens; SKILL.md has 629 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~115
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ankit/stylebot at commit 96a61f5, republished under its MIT licence (© ankit). 629 words, ~1,096 tokens.

Download SKILL.mdSave it as .claude/skills/eval-chat/SKILL.md (or your agent's skills folder).
name
eval-chat
description
Measure whether a change makes Chat's styling replies better, with `yarn eval:chat` — it runs styling requests on recorded real pages for two versions of Stylebot and measures on the page whether each request was met. Use after changing Chat's prompt, the page outline or CSS context it sends, or how replies are applied or checked, before claiming the change helps; when the user asks "is this better?", "run the evals", or wants numbers for a PR.

Evaluate Chat's styling

Unit tests show Chat's plumbing works, not that its replies look better. yarn eval:chat measures that: each case is a request on a recorded page ("Make this page an everforest theme with Fira Code as typography" on Hacker News), run for a base and a head version, then measured on the page (a column's width, a sidebar hidden, a font, a color) and shown in screenshots. Model calls go through headless Claude Code on the user's subscription, not an API key. docs/chat-evals.md describes the harness; this is how to use it well.

1. Run it

From the checkout with the change, after yarn install:

yarn eval:chat --base HEAD --head .
  • Compare against the previous commit (--base HEAD --head .) while iterating, so the numbers isolate this change; compare against v4 for a PR's description.
  • Iterate on all six cases, or the tags a change targets (theme, readability, typography, layout, hide, detailed, variables), one run. Before a PR, run them with --runs 3. To compare models or thinking, give head its own setup with --head-model or --head-thinking.
  • Run it in the background and wait for Results in …; it ends by saying how many model calls it made and how long it took. Results are cached by the code that produced them, so an unchanged base costs nothing the second time.
  • New pages are recorded on first use. Check a new site loads headless before adding it: Reddit and Stack Overflow block it, and some news sites cover the page with a consent dialog that the screenshots and checks then measure instead of the page.

2. Read it

Open summary.md in the results folder.

  • Read "Case by case": every check for base and head, with the value the page had when one failed, then the original page and each version's result. Checks are the score; the screenshots are for what no check measures, like whether a theme looks good.
  • When reporting a run, show it: each case's checks for base and head, and its screenshots, not only the overall numbers. One run is a hint, three are evidence.
  • Look before you believe a surprising score. Each case's folder has the screenshots, the conversation and the stylesheet. Harness bugs have looked like model failures: web fonts not loaded before the screenshot, a page's CSP blocking the injected script.
  • Look for failures that repeat across cases; those are what to fix. When the screenshots show a problem no check catches, add a check (see docs/chat-evals.md).
  • Changing a case's checks? Run yarn eval:chat --references --cases <id> first: no model calls, and every check should pass with the case's reference stylesheet. Write a reference for a case that doesn't have one yet.
Show full SKILL.md (187 more words)Show less

3. Improve what it finds

What moved the scores, roughly in order:

  • Context before wording. The biggest wins came from what the model was shown, not from instructions: an outline cut off before the footer, a header with no selectable trait until the outline named its bgcolor, a design system's base palette buried under thousands of CSS variables.
  • Give absolute values where the model can't see the current ones. It can't see spacing, so "make it compact" nudged padding and made rows taller.
  • Feedback on the model's own edits must name the cause. Naming the symptom made a fix recolor every title dark on dark.
  • Watch for instructions that conflict on real pages: "links get the accent" turned list sites, whose titles are links, into a wall of one color.
  • Re-run the cases a change might hurt, not only the ones it targets: telling the model to be bold on look requests helped those and hurt others.

4. Report it

In the PR description, give the comparison against v4: the overall row (checks, unreadable, zero-match, cost), the model setup and runs, and any case that got worse.

© ankit, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/eval-chat of ankit/stylebot.

Open the folder on GitHubat commit 96a61f5

Compare with similar skills

Eval Chat next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Chat compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Chat this skillankit/stylebot1.6k—~1.1kAutomated safety check: PassMIT
Tui Validatemikeyobrien/ralph-orchestrator3.2k—~3kAutomated safety check: PassMIT
LLM Trace Review Interfaceai-evals-course/evals-skills1.5k—~1.4kAutomated safety check: PassApache-2.0
EvaluationPrimeIntellect-ai/prime-envs130—~4.6kAutomated safety check: PassApache-2.0
Eval Guidemicrosoft/eval-guide138—~22kAutomated safety check: WarnMIT
Write A Specdifferent-ai/openwork24k—~3.3kAutomated safety check: PassCustom licence

Similar skills

  • Tui Validate

    mikeyobrien/ralph-orchestrator

    Validates Terminal User Interface (TUI) output using freeze for screenshot capture and LLM-as-judge for semantic validation.

    3.2k GitHub stars~3k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Trace Review Interface

    ai-evals-course/evals-skills

    Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data.

    1.5k GitHub stars~1.4k tokensUpdated 14 days ago
    AI & LLM EngineeringAuto-check passed
  • Evaluation

    PrimeIntellect-ai/prime-envs

    Install and run a verifiers environment — smoke testing during development and full benchmark evals.

    130 GitHub stars~4.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Eval Guide

    microsoft/eval-guide

    Official

    Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately.

    138 GitHub stars~22k tokensUpdated 3 mo ago
    Testing & QAAuto-check: warnings
  • Write A Spec

    different-ai/openwork

    Write or extend an E2E journey spec in evals/specs that proves a PR's change to a human reviewer.

    24k GitHub stars~3.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Caveman Experiment Manager

    JuliusBrussee/caveman

    Reads the state and results of Caveman Cloud experiments and reports one recommendation or a block, without changing an experiment's lifecycle itself.

    110k GitHub starsUsed in 1 repo~975 tokens
    AI & LLM EngineeringAuto-check passed

More from ankit/stylebot

  • Release

    ankit/stylebot

    Cut a Stylebot release — CHANGELOG entry, version bump, release PR, tag, and store zips.

    1.6k GitHub stars~897 tokensUpdated today
    Auto-check passed
  • Translate

    ankit/stylebot

    Translate Stylebot's locale strings — fill in strings that English has and other locales lack, retranslate English strings that changed, or add a whole new locale (e.g.

    1.6k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Validate Stylebot

    ankit/stylebot

    Launch Stylebot in a real, headed Chrome window (via yarn dev:chrome) so the user can manually eyeball a change — in whatever checkout the current session is in, worktree or main.

    1.6k GitHub stars~780 tokensUpdated today
    Auto-check passed
  • Stylebot

    ankit/stylebot

    Restyle websites in the user's browser with the Stylebot extension's stylebot CLI — dark modes, themes, fonts, bigger text, hiding clutter, fixing a layout, or editing and managing the CSS and…

    1.6k GitHub stars~2.8k tokensUpdated today
    Auto-check passed

Questions about Eval Chat

What does Eval Chat do?

Measure whether a change makes Chat's styling replies better, with yarn eval:chat — it runs styling requests on recorded real pages for two versions of Stylebot and measures on the page whether each…. Eval Chat is an agent skill from ankit/stylebot. Measure whether a change makes Chat's styling replies better, with yarn eval:chat — it runs styling requests on recorded real pages for two versions of Stylebot and measures on the page whether each request was met.

When should I use Eval Chat?

Eval Chat fits situations like: asks is this better?; wants numbers for a PR.

How do I install Eval Chat in Claude Code?

Run `npx skills add ankit/stylebot --skill eval-chat -a claude-code`. Or copy the skill folder (.agents/skills/eval-chat in ankit/stylebot) into .claude/skills/eval-chat in your project. Claude Code loads it when a task matches its description.

How do I install Eval Chat in Codex?

Run `npx skills add ankit/stylebot --skill eval-chat -a codex`. Or copy the skill folder (.agents/skills/eval-chat in ankit/stylebot) into .agents/skills/eval-chat in your project. Codex loads it when a task matches its description.

Can I use Eval Chat in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ankit/stylebot --skill eval-chat -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-chat, .gemini/skills/eval-chat, .github/skills/eval-chat and .opencode/skills/eval-chat in your project.

What does Eval Chat need to run?

Going by SKILL.md and its folder, Eval Chat needs the command-line tools its instructions call (yarn).

Does Eval Chat access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Chat safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Chat use?

Eval Chat is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Chat use?

About 1.1k tokens (SKILL.md is roughly 4.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Chat?

Skills that share tags, products or a category with Eval Chat: Tui Validate (mikeyobrien/ralph-orchestrator, 3.2k stars), LLM Trace Review Interface (ai-evals-course/evals-skills, 1.5k stars), Evaluation (PrimeIntellect-ai/prime-envs, 130 stars) and Eval Guide (microsoft/eval-guide, 138 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Chat?

ankit (a GitHub user) maintains it in ankit/stylebot, which has 1,647 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 8, 2026.

Source: ankit/stylebot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.