Evaluate the quality of any AI-generated artifact — visualizations, code, documents, conversations, or any skill output.

MITAuto-check passedDocuments & Office

Install Evaluate

skills CLI
$ npx skills add careerhackeralex/visualize --skill evaluate -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install careerhackeralex/visualize evaluate --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/careerhackeralex/visualize.git skills-src && mkdir -p .claude/skills && cp -r skills-src/eval .claude/skills/evaluate && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluate
GitHub stars
235
Token cost
~2.4k tokens
SKILL.md length
542 words
Files
621
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Evaluate the quality of any AI-generated artifact — visualizations, code, documents, conversations, or any skill output.

  • Works in 3 steps: Spec Generation → Evaluation → Visual Report (via /visualize)
  • Tasks that involve Slides and decks
  • SKILL.md covers How It Works, Phase 1: Spec Generation, Phase 2: Evaluation and Phase 3: Visual Report (via…, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Evaluate is an agent skill from careerhackeralex/visualize. Evaluate the quality of any AI-generated artifact — visualizations, code, documents, conversations, or any skill output. Works in 3 phases: (1) Generate evaluation specs tailored to the artifact type, (2) Run comprehensive evaluation against those specs, (3) Produce a beautiful visual report using the /visualize skill. Use after any skill produces output, or invoke directly with /evaluate <file-or-context. Supports evaluating: HTML visualizations, code projects, documents, agent conversations, slide decks…

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 621 other files (for example `EVAL.md`, `LOOP.md` and `OPENCLAW-SETUP.md`).

It sits in Documents & Office, covering Slides and decks. The repository describes itself as: Turn any idea into a beautiful HTML visualization — with one prompt. A Claude Code skill. The licence is MIT.

When your agent uses it

  • Tasks that involve Slides and decks

Example prompts

  • “/evaluate”

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Spec Generation
  2. Evaluation
  3. Visual Report (via /visualize)

What it can do on your machine

Read from SKILL.md and the folder at commit 3010770. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are javascript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluate loads about 2.4k tokens when it runs. Until then it costs about 144 tokens; SKILL.md has 542 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~144
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from careerhackeralex/visualize at commit 3010770, republished under its MIT licence (© careerhackeralex). 542 words, ~2,387 tokens.

Download SKILL.mdSave it as .claude/skills/evaluate/SKILL.md (or your agent's skills folder). This skill also uses 620 other files; get the full folder from GitHub.
name
evaluate
description
Evaluate the quality of any AI-generated artifact — visualizations, code, documents, conversations, or any skill output. Works in 3 phases: (1) Generate evaluation specs tailored to the artifact type, (2) Run comprehensive evaluation against those specs, (3) Produce a beautiful visual report using the /visualize skill. Use after any skill produces output, or invoke directly with /evaluate <file-or-context>. Supports evaluating: HTML visualizations, code projects, documents, agent conversations, slide decks, dashboards, or any artifact with quality dimensions.

Evaluate

Comprehensive quality evaluation for any AI-generated artifact. Produces its report as a visualization.

How It Works

┌──────────────────────────────────────────────┐
│                                              │
│  Phase 1: SPEC GENERATION                    │
│  Analyze the artifact type                   │
│  Generate tailored evaluation criteria       │
│  Define scoring dimensions + weights         │
│  Set quality gates                           │
│           │                                  │
│           ▼                                  │
│  Phase 2: EVALUATION                         │
│  Run automated checks (when possible)        │
│  Visual/manual inspection                    │
│  Score each dimension with evidence          │
│  Identify systemic vs local issues           │
│           │                                  │
│           ▼                                  │
│  Phase 3: REPORT (via /visualize)            │
│  Generate a beautiful HTML eval report       │
│  Scores, charts, screenshots, fix list       │
│  Radar chart of dimensions                   │
│  Before/after tracking                       │
│                                              │
└──────────────────────────────────────────────┘

Phase 1: Spec Generation

For any artifact, generate evaluation specs by analyzing:

1. Identify Artifact Type
  • HTML Visualization → visual design, interactivity, technical, content, shareability
  • Code/Project → correctness, readability, architecture, test coverage, performance
  • Document/Report → clarity, structure, accuracy, completeness, tone
  • Conversation/Agent → helpfulness, accuracy, tone, efficiency, safety
  • Slide Deck → all visualization dims + narrative flow, persuasion, pacing
  • Dashboard → data accuracy, information density, scannability, actionability
  • Custom → derive dimensions from the skill's SKILL.md and stated goals
2. Generate Dimensions

For each artifact type, produce 6-10 evaluation dimensions. Each dimension needs:

  • Name — short, clear label
  • Description — what this dimension measures
  • Weight — percentage (all weights sum to 100%)
  • Scoring anchors — what does a 10, 8, 6, 4 look like?
  • Automated checks — any programmatic tests (if applicable)
  • Deductions — specific issues and their point costs
3. Set Quality Gates

Define gates based on the artifact's purpose:

GateCriteriaMeaning
🚀 EXCEPTIONALOverall ≥ 9.5, all ≥ 9Best-in-class. Share everywhere.
✅ SHIPOverall ≥ 9.0, all ≥ 8Production-ready.
⚠️ ACCEPTABLEOverall ≥ 8.0, all ≥ 7Usable but not impressive.
🔧 NEEDS WORKOverall ≥ 7.0 or any < 7Fix before releasing.
❌ FAILOverall < 7.0 or any < 5Major rework.
4. Output Spec Document

Write the spec to eval-spec-[artifact-name].md for reference and reuse.

Phase 2: Evaluation

For HTML Visualizations

Open in browser at 3 viewports (1280×720, 768×1024, 375×667).

Automated audit (run in browser console):

javascript
(function() {
  const audit = {};
  const style = [...document.querySelectorAll('style')].map(s => s.textContent).join(' ');
  const html = document.documentElement.outerHTML;

  // Structure
  audit.hasDoctype = /^<!doctype html>/i.test(html);
  audit.hasLangAttr = !!document.documentElement.lang;
  audit.hasCharset = !!document.querySelector('meta[charset]');
  audit.hasViewport = !!document.querySelector('meta[name="viewport"]');
  audit.hasTitle = document.title.length > 0;

  // Menu system
  audit.menuExists = !!document.querySelector('.viz-menu');
  audit.menuHasTheme = !!html.match(/cycleTheme|themeLabel/i);
  audit.menuHasDownload = !!html.match(/htmlToImage|html-to-image/i);
  audit.menuHasPrint = !!html.match(/window\.print/i);

  // Theme system
  audit.hasCSSVars = !!style.match(/--bg\s*:/);
  audit.hasDarkTheme = !!style.match(/(\.theme-dark|:root)[\s\S]*?--bg/);
  audit.hasLightTheme = !!style.match(/\.theme-light/);
  audit.themePersistedToStorage = !!html.match(/localStorage.*theme/i);

  // Typography
  audit.hasInterFont = !!html.match(/fonts\.googleapis.*Inter|font-family.*Inter/i);
  audit.hasFontFallback = !!style.match(/-apple-system|system-ui/);
  audit.bodyFontSize = parseFloat(getComputedStyle(document.body).fontSize);
  audit.bodyFontOK = audit.bodyFontSize >= 14;

  // Layout
  audit.usesFlexOrGrid = !!(style.match(/display\s*:\s*(flex|grid)/));
  audit.hasMaxWidth = !!style.match(/max-width/);
  audit.hasResponsiveBreakpoints = !!style.match(/@media.*max-width|@media.*min-width|sm:|md:|lg:/);

  // Print & Accessibility
  audit.hasPrintStyles = !!style.match(/@media\s*print/);
  audit.hasPrintColorAdjust = !!style.match(/print-color-adjust/);
  audit.hasReducedMotion = !!style.match(/prefers-reduced-motion/);
  audit.hasAriaLabels = !!html.match(/aria-label/);
  audit.hasSemanticHTML = !!html.match(/<(header|main|nav|section|article|footer)/);

  // Animations
  audit.hasKeyframes = !!style.match(/@keyframes/);
  audit.hasTransitions = !!style.match(/transition\s*:/);

  // Performance
  audit.fileSizeKB = Math.round(new Blob([html]).size / 1024);
  audit.fileSizeOK = audit.fileSizeKB < 200;
  audit.noExternalImages = document.querySelectorAll('img[src^="http"]').length === 0;
  audit.htmlToImageLoaded = typeof htmlToImage !== 'undefined';

  // Summary
  const bools = Object.entries(audit).filter(([k,v]) => typeof v === 'boolean');
  const passed = bools.filter(([k,v]) => v).length;
  audit._passed = passed;
  audit._total = bools.length;
  audit._percent = Math.round(passed / bools.length * 100);
  audit._failures = bools.filter(([k,v]) => !v).map(([k]) => k);

  console.table(audit);
  return audit;
})();

Visual scoring — 8 dimensions for visualizations:

#DimensionWeight10 =6 =
D1First Impression15%Apple keynote qualityGeneric template feel
D2Typography15%Perfect hierarchy, Inter font, fluid sizingAll same size, no hierarchy
D3Color & Contrast10%Harmonious, WCAG AA, both themes beautifulClashing, low contrast
D4Layout & Spacing15%Consistent rhythm, responsive, generous spaceCramped, broken at mobile
D5Content Quality15%Clear message in 5 seconds, zero fillerConfusing, placeholder text
D6Interactivity10%Menu + theme + download + print all flawlessMissing features, broken
D7Technical10%Zero errors, semantic, accessible, print-readyConsole errors, broken layout
D8Shareability10%Would tweet this unpromptedWorse than Canva
Show full SKILL.md (204 more words)Show less
For Code/Projects

Dimensions: Correctness, Readability, Architecture, Error Handling, Performance, Testing, Documentation, Security

For Documents

Dimensions: Clarity, Structure, Accuracy, Completeness, Tone, Formatting, Actionability, Brevity

For Agent Conversations

Dimensions: Helpfulness, Accuracy, Tone, Efficiency, Safety, Context Awareness, Tool Usage, Follow-through

Phase 3: Visual Report (via /visualize)

After scoring, generate the eval report as a beautiful HTML dashboard using the visualize skill:

Report Structure
  1. Hero — artifact name, overall score (big number), quality gate badge
  2. Radar Chart — all dimensions plotted on a radar/spider chart (Chart.js)
  3. Dimension Cards — each dimension as a card with score, bar, key notes
  4. Automated Audit — pass/fail checklist with percentages
  5. Screenshots — key views embedded (if HTML artifact)
  6. Fix List — prioritized fixes as a kanban-style layout (critical / high / medium / low)
  7. Systemic Issues — patterns that affect all outputs (flagged for SKILL.md fixes)
  8. History — if re-evaluating, show before/after score comparison chart
Report Filename

eval-report-[artifact-name]-[date].html

The report itself must score ≥ 9.0 on the visualize eval criteria.

This is the ultimate dogfood test — our evaluation tool produces evaluations using our visualization tool.

The Improvement Loop

Generate artifact (any skill)
       ↓
/evaluate → Spec + Score + Visual Report
       ↓
Review report → identify fixes
       ↓
Fix (systemic → SKILL.md, local → artifact)
       ↓
/evaluate again → compare scores
       ↓
Ship when gate = SHIP or EXCEPTIONAL

Max 3 loops per artifact. If it can't reach SHIP in 3 loops, the problem is in the skill — update the skill's instructions, not the artifact.

Quick Start

# Evaluate a visualization
/evaluate path/to/visualization.html

# Evaluate with custom context
/evaluate path/to/code-project --type code

# Re-evaluate after fixes (tracks improvement)
/evaluate path/to/visualization.html --loop 2

# Generate specs only (no scoring)
/evaluate --specs-only --type dashboard

© careerhackeralex, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 620 other files in eval of careerhackeralex/visualize.

  • SKILL.md
  • EVAL.md
  • LOOP.md
  • OPENCLAW-SETUP.md
  • archive/eval-results-raw.md
  • archive/eval-round10-raw.md
  • archive/eval-round11-raw.md
  • archive/eval-round12-raw.md
  • archive/eval-round13-raw.md
  • archive/eval-round14-raw.md
  • archive/eval-round15-raw.md
  • archive/eval-round16-raw.md
  • archive/eval-round17-research.md
  • archive/eval-round18-raw.md
  • archive/eval-round19-raw.md
  • archive/eval-round2-raw.md
  • archive/eval-round20-raw.md
  • archive/eval-round21-raw.md
  • archive/eval-round22-raw.md
  • archive/eval-round23-raw.md
  • … and 601 more

Open the folder on GitHubat commit 3010770

Compare with similar skills

Evaluate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluate compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluate this skillcareerhackeralex/visualize235—~2.4kAutomated safety check: PassMIT
Image To Editable Pptningzimu/image-to-editable-ppt-skill2.8k—~4.3kAutomated safety check: PassMIT
Slidesfcakyon/claude-codex-settings1.2k1 repos~1.1kAutomated safety check: PassMIT
Ppt Image FirstNyxTides/ppt-image-first1.2k—~1.6kAutomated safety check: PassApache-2.0
Vibe to Agentic Engineering Frameworkshanraisshan/claude-code-best-practice67k—~3.3kAutomated safety check: PassMIT
Gpt Image2 PptJuneYaooo/gpt-image2-ppt-skills1.3k—~8.9kAutomated safety check: NotesApache-2.0

Similar skills

  • Image To Editable Ppt

    ningzimu/image-to-editable-ppt-skill

    Rebuild slide images, scanned or image-based PPT/PPTX files, and PDF decks into object-level editable PowerPoint (.pptx), preserving speaker notes when supplied.

    2.8k GitHub stars~4.3k tokensUpdated 21 days ago
    Documents & OfficeAuto-check passed
  • Slides

    fcakyon/claude-codex-settings

    Create and edit presentation slide decks (.pptx) with PptxGenJS, bundled layout helpers, and render/validation utilities.

    1.2k GitHub starsUsed in 1 repo~1.1k tokens
    Documents & OfficeAuto-check passed
  • Ppt Image First

    NyxTides/ppt-image-first

    Build presentation plans for PPT / slides / decks through a conversation-first workflow, then propose multiple visual directions with preview images before writing deck specs.

    1.2k GitHub stars~1.6k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed
  • Vibe to Agentic Engineering Framework

    shanraisshan/claude-code-best-practice

    Explains the conceptual model behind a presentation on moving from unstructured vibe coding to fully configured agentic engineering, including its 4-level scoring system and slide conventions.

    67k GitHub stars~3.3k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Gpt Image2 Ppt

    JuneYaooo/gpt-image2-ppt-skills

    Generate visually striking PPT slides via OpenAI's gpt-image-2 -- use any style in styles/<collection/STYLEID.md or mimic a user-supplied .pptx template; outputs high-res slide PNGs and a 16:9 .pptx.

    1.3k GitHub stars~8.9k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check: notes
  • Compiledeck

    scunning1975/MixtapeTools

    Create and compile beautiful Beamer presentations following the Rhetoric of Decks philosophy.

    469 GitHub starsUsed in 2 repos~3.1k tokens
    Documents & OfficeAuto-check passed

Questions about Evaluate

What does Evaluate do?

Evaluate the quality of any AI-generated artifact — visualizations, code, documents, conversations, or any skill output. Evaluate is an agent skill from careerhackeralex/visualize. Evaluate the quality of any AI-generated artifact — visualizations, code, documents, conversations, or any skill output.

When should I use Evaluate?

Evaluate fits situations like: tasks that involve Slides and decks.

How do I install Evaluate in Claude Code?

Run `npx skills add careerhackeralex/visualize --skill evaluate -a claude-code`. Or copy the skill folder (eval in careerhackeralex/visualize) into .claude/skills/evaluate in your project. Claude Code loads it when a task matches its description.

How do I install Evaluate in Codex?

Run `npx skills add careerhackeralex/visualize --skill evaluate -a codex`. Or copy the skill folder (eval in careerhackeralex/visualize) into .agents/skills/evaluate in your project. Codex loads it when a task matches its description.

Can I use Evaluate in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add careerhackeralex/visualize --skill evaluate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluate, .gemini/skills/evaluate, .github/skills/evaluate and .opencode/skills/evaluate in your project.

What does Evaluate need to run?

SKILL.md names no scripts, command-line tools or credentials: Evaluate is instructions for the agent only.

Does Evaluate access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evaluate safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evaluate use?

Evaluate is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluate use?

About 2.4k tokens (SKILL.md is roughly 9.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evaluate?

Skills that share tags, products or a category with Evaluate: Image To Editable Ppt (ningzimu/image-to-editable-ppt-skill, 2.8k stars), Slides (fcakyon/claude-codex-settings, 1.2k stars), Ppt Image First (NyxTides/ppt-image-first, 1.2k stars) and Vibe to Agentic Engineering Framework (shanraisshan/claude-code-best-practice, 67k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluate?

careerhackeralex (a GitHub user) maintains it in careerhackeralex/visualize, which has 235 GitHub stars. The repository was last updated on March 6, 2026.

Source: careerhackeralex/visualize on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.