Agent skill

Eval Graphics

by elvisun in elvisun/newsjack

Turn an eval study's numbers into on-brand, publish-ready figures using the Newsjack chart room (the eval design system), then validate them with Playwright.

MITAuto-check passedFrontend & Design

Install Eval Graphics

skills CLI
$ npx skills add elvisun/newsjack --skill eval-graphics -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install elvisun/newsjack eval-graphics --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/elvisun/newsjack.git skills-src && mkdir -p .claude/skills && cp -r skills-src/eval/design-system .claude/skills/eval-graphics && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-graphics
GitHub stars
1.5k
Token cost
~2k tokens
SKILL.md length
985 words
Files
8 (incl. scripts, assets)
Skills in repo
30
Repo updated
First seen
Licence
MIT

At a glance

Turn an eval study's numbers into on-brand, publish-ready figures using the Newsjack chart room (the eval design system), then validate them with Playwright.

  • Works in 6 steps: Read the study's numbers (e.g. an… → Scaffold an HTML file in the study's run… → Lift the matching figure block(s) from… → …
  • Tasks that involve Design systems
  • SKILL.md covers Files you work with, The grammar (non-negotiable…, Pick the figure to fit the data and How to compute geometry…, plus 3 more sections
  • Runs JavaScript scripts from its folder; calls node

What it does

Eval Graphics is an agent skill from elvisun/newsjack. Turn an eval study's numbers into on-brand, publish-ready figures using the Newsjack chart room (the eval design system), then validate them with Playwright. For producing the charts in a published eval/data study.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and assets (for example `README.md` and `scripts/package.json`).

It sits in Frontend & Design, covering Design systems and Browser testing. It works with Playwright. The repository describes itself as: The open-source skills that turn your agent into a full PR team. The licence is MIT.

When your agent uses it

  • Tasks that involve Design systems
  • Tasks that involve Browser testing

Example prompts

  • “/eval-graphics”

Requirements

  • Node.js

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Read the study's numbers (e.g. an aggregate.py printout or
  2. Scaffold an HTML file in the study's run folder, e.g.
  3. Lift the matching figure block(s) from chart-room.html. Recompute every
  4. Keep it honest. Label sample sizes (n=), say what the score is, don't
  5. Validate with Playwright (required)
  6. Eyeball the screenshot before publishing. Read the full-page PNG; confirm

What it can do on your machine

Read from SKILL.md and the folder at commit b5a8dc8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Graphics loads about 2k tokens when it runs. Until then it costs about 57 tokens; SKILL.md has 985 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~57
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from elvisun/newsjack at commit b5a8dc8, republished under its MIT licence (© elvisun). 985 words, ~2,027 tokens.

Download SKILL.mdSave it as .claude/skills/eval-graphics/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
eval-graphics
description
Turn an eval study's numbers into on-brand, publish-ready figures using the Newsjack chart room (the eval design system), then validate them with Playwright. For producing the charts in a published eval/data study.
scope
eval-internal
not_a_product_skill
true
when_to_use
You have an eval study's results (a results.json, an aggregate table, head-to-head numbers, per-dimension deltas) and need branded figures to publish it — a…

Eval Graphics — the chart room

Internal tooling, not a product skill. This lives in eval/design-system/ and is used by maintainers to publish eval studies. It is not a Newsjack user skill, must never be installed into skills/, and is never loaded at product runtime. Do not confuse it with the main skills folder.

You produce figures for an eval study: standalone HTML that renders the Newsjack house chart style, validated by Playwright and screenshotted to PNGs you drop into a writeup. One grammar — newsprint paper, ink, a single vermilion mark — across every figure, so a study reads like one publication.

Files you work with

  • assets/colors_and_type.css — design tokens (palette, type, spacing). Never edit; always link.
  • charts.css — the chart primitives (.bar-primary, .line-base, .fig, .stat, heatmap classes, masthead/section/colophon scaffold). Never edit; always link.
  • chart-room.html — the specimen gallery of all 9 figure types with placeholder data. This is your copy-paste source: find the figure that fits your data, lift its block, swap the geometry and labels.
  • scripts/validate.mjs — Playwright validator + screenshotter.

The grammar (non-negotiable brand rules)

These come straight from the design system. Breaking one means the figure is off-brand:

  • One chroma. Vermilion #E05A47 is the only colour. The highlighted / winning series is accent; every other series is quiet grey (--c-base, #E4E0D8) or an ink wash. If you reach for a second colour, you've gone wrong. No gradients. No emoji. Ever.
  • Paper, not white. Background is --nj-page #F9F8F6. Borders are the hairline rgba(26,26,26,0.10). Cards are 0-radius with a faint editorial shadow.
  • Type roles. Figure titles: Newsreader italic. Axis ticks, value labels, legends, eyebrows: IBM Plex Mono, ALL CAPS, ≥0.12em tracking. Descriptive captions: DM Sans.
  • Value labels float above the bar (mono, centered over the column).
  • Headline figures are one consistent colour — the whole numeral and its symbol in the accent (e.g. +38%, 4.2× fully vermilion), via <span class="accent">.
  • Section pattern: top hairline → mono number + lowercase-italic title (left)
    • mono subtitle (right) → content.
Accent strategy — set on <body>
data-accentUse
winner (default)best/highlighted series = accent, others grey. The standard comparison look.
singlecomparisons go all-grey; accent is reserved for one hero mark (a donut, one bar).
monoeverything ink, no chroma — for a sober, neutral data study.

Also data-grid="on|off" (gridlines) and data-barstyle="solid|outline".

Pick the figure to fit the data

Data shapeFigure (block in chart-room.html)
Two series across a few benchmarksFIG.01 grouped bars — the flagship
One series, ranked by categoryFIG.02 single-series bars
Before → after on sparse metricsFIG.03 dumbbell / lollipop
A value over time / scale / versionsFIG.04 line / scaling curve
Two-axis tradeoff, one point highlightedFIG.05 scatter / quadrant
Composition / share across rowsFIG.06 100% stacked bars
One number that deserves the frameFIG.07 donut
Model × task (or any) matrixFIG.08 heatmap (data-driven JS)
2–4 punchy single numbers, no axesFIG.09 big-stat callouts

How to compute geometry (filling an SVG template)

The SVGs use plain coordinates inside a viewBox. The mapping math you need:

Vertical bars (FIG.01/02). Choose a baseline yBase (value 0) and a top yTop (max value). scale = (yBase - yTop) / maxVal. For a value v: barHeight = v * scale, barY = yBase - barHeight, value label at y = barY - 12. (Specimen FIG.01: yBase=500, 100→y=120, so scale=3.8.)

Dumbbell (FIG.03). Horizontal axis from xMin (value 0) to xMax (value 100). x(v) = xMin + (v/ (maxVal)) * (xMax - xMin). Draw a .stem line from x(before) to x(after), a .dot-base at before, a .dot-primary at after.

Line (FIG.04). y(v) = yBase - v*scale; evenly space x across the points; .line-primary for focal, .line-base (dashed) for baseline; optional .area-primary polygon closes down to yBase.

Donut (FIG.07). C = 2 * π * r (specimen r=104 → C≈653.45). For percent p: dash = (p/100) * C; set stroke-dasharray="{dash} {C}" on the accent ring, transform="rotate(-90 cx cy)" so it starts at 12 o'clock.

Heatmap (FIG.08). Don't hand-place cells — edit the data, tasks, models arrays in the inline <script>. Cell alpha = 0.12 + clamp((v-min)/(max-min)) * 0.88; text flips to white above alpha 0.55. Already handles any grid size.

Big-stat (FIG.09). Pure markup: <span class="fig-num"><span class="accent"> +38</span>%</span> — wrap the whole figure (or the part that should be coloured) in .accent. Keep the numeral and its symbol the same colour.

Show full SKILL.md (322 more words)Show less

Build process

  1. Read the study's numbers (e.g. an aggregate.py printout or results.json). Decide which 1–4 figures tell the story; don't over-chart.
  2. Scaffold an HTML file in the study's run folder, e.g. eval/<study>/runs/<run>/figures/<name>.html. Link the design system by relative path: <link rel="stylesheet" href="../../../../design-system/assets/colors_and_type.css"> and .../design-system/charts.css. (Count the ../ from the figure file to eval/design-system/.) Set <body data-accent="…">.
  3. Lift the matching figure block(s) from chart-room.html. Recompute every coordinate from the real data using the math above. Replace category labels, value labels, series names, captions, the masthead/section text. Delete the figures you don't use.
  4. Keep it honest. Label sample sizes (n=), say what the score is, don't round a tie into a win, and don't invent a series the data doesn't have. The angle-generator/meanest-editor anti-slop ethos applies to charts too.
  5. Validate with Playwright (required):
    bash
    cd eval/design-system/scripts
    node validate.mjs ../../<study>/runs/<run>/figures/<name>.html --out ../../<study>/runs/<run>/figures/png
    All checks must pass (no JS errors, accent painted, serif applied, every figure has geometry, no empty SVGs). It writes a full-page PNG plus one crop per figure — those are your embeddable assets.
  6. Eyeball the screenshot before publishing. Read the full-page PNG; confirm bars/labels line up and nothing overflows.

Worked example

For the Fable-5-vs-Opus-4.8 study (eval/fable-vs-opus/), the headline numbers (overall dim mean 4.60 vs 4.36; 67/33 head-to-head; 24 vs 7 robust wins; per- dimension deltas) map cleanly to: a FIG.01 grouped-bars of the 7 dimensions (Fable accent, Opus grey), a FIG.07 donut for the win rate, and a row of FIG.09 big-stat callouts (overall mean, robust wins, publishable rate). See eval/fable-vs-opus/runs/2026-06-09-full/figures/ for the built, validated set.

Guardrails

  • Never edit assets/colors_and_type.css or charts.css to force a one-off look — if a figure needs something the system lacks, that's a design-system change, raise it, don't hack the study.
  • Never add a second colour, a gradient, a drop shadow on a chart, or emoji.
  • Never ship a figure that hasn't passed validate.mjs.
  • Placeholder data stays clearly labelled until real numbers replace it.

© elvisun, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, assets) in eval/design-system of elvisun/newsjack.

  • SKILL.md
  • .gitignore
  • README.md
  • assets/colors_and_type.css
  • chart-room.html
  • charts.css
  • scripts/package.json
  • scripts/validate.mjs

Open the folder on GitHubat commit b5a8dc8

Compare with similar skills

Eval Graphics next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Graphics compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Graphics this skillelvisun/newsjack1.5k—~2kAutomated safety check: PassMIT
Igniteui Angular Figma To AppIgniteUI/igniteui-angular599—~6.4kAutomated safety check: PassMIT
Extract Design Systemespennilsen/pi122—~2kAutomated safety check: PassMIT
Frontend Visual QAdaymade/claude-code-skills1.4k—~7.7kAutomated safety check: PassMIT
Reprovaadin/web-components582—~1.3kAutomated safety check: PassNone
Scout UI Testingelastic/kibana21k—~3.1kAutomated safety check: PassCustom licence

Similar skills

  • Igniteui Angular Figma To App

    IgniteUI/igniteui-angular

    Builds Angular views from Figma designs with Ignite UI for Angular, supporting Indigo.Design kits, third-party kits, and plain frames.

    599 GitHub stars~6.4k tokensUpdated today
    Frontend & DesignAuto-check passed
  • Extract Design System

    espennilsen/pi

    Reverse-engineer a design system from a live website (public URL or localhost).

    122 GitHub stars~2k tokensUpdated 16 days ago
    Frontend & DesignAuto-check passed
  • Frontend Visual QA

    daymade/claude-code-skills

    Audits already-rendered UI — web app, deck/slide, dashboard, design-system, or Electron/native app — via real-browser/native-app journeys and a Playwright sweep.

    1.4k GitHub stars~7.7k tokensUpdated today
    Frontend & DesignAuto-check passed
  • Repro

    vaadin/web-components

    Reproduce a Vaadin web component bug from a GitHub issue in vaadin/web-components.

    582 GitHub stars~1.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Scout UI Testing

    elastic/kibana

    Official

    A skill your agent uses when creating, updating, debugging, or reviewing Scout UI tests in Kibana (Playwright + Scout fixtures), including page objects, browser authentication, parallel UI tests…

    21k GitHub stars~3.1k tokensUpdated today
    Frontend & DesignAuto-check passed
  • Liteyuki Webui Frontend

    LiteyukiStudio/LiteyukiBot

    Build, review, or test LiteyukiBot v7's React/Vite WebUI under webui/ and its packaged static delivery in packages/webui/.

    157 GitHub stars~2.5k tokensUpdated 1 mo ago
    Frontend & DesignAuto-check passed

More from elvisun/newsjack

All 30 skills in this repo
  • AI Visibility Writing

    elvisun/newsjack

    Audit, question, suggest, or fact-preservingly revise a press release, blog post, contributed article, or expert explainer so AI answer systems can more easily retrieve, understand, quote, and cite…

    1.5k GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Icp Evidence Analysis

    elvisun/newsjack

    Turn a source-bound company and market dossier into testable ideal-customer-profile hypotheses, buying roles, triggers, constraints, disqualifiers, standing, counterevidence, and research gaps.

    1.5k GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Research any company, product, or service from a URL plus description and build a comprehensive, evidence-bound AEO/GEO/AI-visibility prompt panel across buyer jobs, information acts, journey…

    1.5k GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Reactive Comment

    elvisun/newsjack

    Triage inbound journalist source queries and draft a response only when the user's expertise is a real fit.

    1.5k GitHub stars~5.6k tokensUpdated yesterday
    Auto-check passed
  • Recover source-bound buyer jobs, struggling moments, desired progress, forces, workarounds, information acts, journey states, criteria, constraints, roles, locales, and authentic language.

    1.5k GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Coverage Tracker

    elvisun/newsjack

    Run a Google Alerts-style keyword coverage tracker. An agent skill from elvisun/newsjack.

    1.5k GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Eval Graphics

What does Eval Graphics do?

Turn an eval study's numbers into on-brand, publish-ready figures using the Newsjack chart room (the eval design system), then validate them with Playwright. Eval Graphics is an agent skill from elvisun/newsjack. Turn an eval study's numbers into on-brand, publish-ready figures using the Newsjack chart room (the eval design system), then validate them with Playwright.

When should I use Eval Graphics?

Eval Graphics fits situations like: tasks that involve Design systems; tasks that involve Browser testing.

How do I install Eval Graphics in Claude Code?

Run `npx skills add elvisun/newsjack --skill eval-graphics -a claude-code`. Or copy the skill folder (eval/design-system in elvisun/newsjack) into .claude/skills/eval-graphics in your project. Claude Code loads it when a task matches its description.

How do I install Eval Graphics in Codex?

Run `npx skills add elvisun/newsjack --skill eval-graphics -a codex`. Or copy the skill folder (eval/design-system in elvisun/newsjack) into .agents/skills/eval-graphics in your project. Codex loads it when a task matches its description.

Can I use Eval Graphics in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add elvisun/newsjack --skill eval-graphics -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-graphics, .gemini/skills/eval-graphics, .github/skills/eval-graphics and .opencode/skills/eval-graphics in your project.

What does Eval Graphics need to run?

Going by SKILL.md and its folder, Eval Graphics needs JavaScript for the scripts in its folder and the command-line tools its instructions call (node). Our summary lists: Node.js.

Does Eval Graphics access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Graphics safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Eval Graphics use?

Eval Graphics is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Graphics use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Graphics?

Skills that share tags, products or a category with Eval Graphics: Igniteui Angular Figma To App (IgniteUI/igniteui-angular, 599 stars), Extract Design System (espennilsen/pi, 122 stars), Frontend Visual QA (daymade/claude-code-skills, 1.4k stars) and Repro (vaadin/web-components, 582 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Graphics?

elvisun (a GitHub user) maintains it in elvisun/newsjack, which has 1,520 GitHub stars. The repository holds 30 skills in this directory. The repository was last updated on October 7, 2026.

Source: elvisun/newsjack on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.