Agent skill

Heuristic Evaluation

by Owl-Listener in Owl-Listener/designpowers

Use after a build to evaluate usability against Nielsen's 10 heuristics and run cognitive walkthroughs of every key task.

MITAuto-check passedFrontend & Design

Install Heuristic Evaluation

skills CLI
$ npx skills add Owl-Listener/designpowers --skill heuristic-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Owl-Listener/designpowers heuristic-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Owl-Listener/designpowers.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/heuristic-evaluation .claude/skills/heuristic-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
heuristic-evaluation
GitHub stars
251
Token cost
~1.3k tokens
SKILL.md length
604 words
Files
1
Skills in repo
33
Repo updated
First seen
Licence
MIT

At a glance

Use after a build to evaluate usability against Nielsen's 10 heuristics and run cognitive walkthroughs of every key task.

  • Works in 5 steps: Confirm Inputs Exist → Dispatch the heuristic-evaluator Agent → Run in Parallel and Reconcile → …
  • Tasks that involve UX design
  • SKILL.md covers When to Use, Process, What You Get and Integration
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Heuristic Evaluation is an agent skill from Owl-Listener/designpowers. Use after a build to evaluate usability against Nielsen's 10 heuristics and run cognitive walkthroughs of every key task. Dispatches the heuristic-evaluator agent in parallel with design-critic and accessibility-reviewer, then feeds findings into the reconciliation protocol. This is the usability lens — "will people actually be able to use this?" — distinct from craft critique and accessibility audit

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Frontend & Design, covering UX design and Accessibility. The repository describes itself as: An agent design team you control: 10 agents that run an inclusive design process while you direct. The licence is MIT.

When your agent uses it

  • Tasks that involve UX design
  • Tasks that involve Accessibility

Example prompts

  • “will people actually be able to use this?”
  • “/heuristic-evaluation”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Confirm Inputs Exist
  2. Dispatch the heuristic-evaluator Agent
  3. Run in Parallel and Reconcile
  4. Honour the Auto-Mode Safeguard
  5. Feed the Fix Round

What it can do on your machine

Read from SKILL.md and the folder at commit cb00757. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Heuristic Evaluation loads about 1.3k tokens when it runs. Until then it costs about 106 tokens; SKILL.md has 604 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~106
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Owl-Listener/designpowers at commit cb00757, republished under its MIT licence (© Owl-Listener). 604 words, ~1,292 tokens.

Download SKILL.mdSave it as .claude/skills/heuristic-evaluation/SKILL.md (or your agent's skills folder).
name
heuristic-evaluation
description
Use after a build to evaluate usability against Nielsen's 10 heuristics and run cognitive walkthroughs of every key task. Dispatches the heuristic-evaluator agent in parallel with design-critic and accessibility-reviewer, then feeds findings into the reconciliation protocol. This is the usability lens — "will people actually be able to use this?" — distinct from craft critique and accessibility audit

Heuristic Evaluation

Heuristic evaluation is the usability lens on a build. Where designpowers-critique asks "does this match the plan?" and the accessibility review asks "can everyone access this?", heuristic evaluation asks "will people actually be able to use this without getting lost, confused, or stuck?" This skill dispatches the heuristic-evaluator agent and integrates its findings with the other reviewers.

When to Use

  • After design-builder completes a build — run alongside designpowers-critique and the accessibility review, not after them
  • Before a fix round, so usability findings are reconciled with craft and accessibility findings together
  • When a flow feels confusing but you can't articulate why — the heuristics name the problem
  • When evaluating an existing design (see using-designpowers → Review Mode), not just freshly built work

This skill is not optional after a build. Skipping it is a Red Flag (see using-designpowers). Usability problems that craft critique and accessibility audits both miss are exactly what this catches.

Process

Step 1: Confirm Inputs Exist

Before dispatching, assemble:

  • The build to evaluate (running app, prototype, or screenshots — test what was built, not the spec)
  • The design brief (from design-discovery) — for the key tasks to walk through
  • The personas (from inclusive-personas) — each persona's primary task gets a cognitive walkthrough
  • design-state.md — the shared state the agent will read and update

If there is no brief or no list of key tasks, the cognitive walkthrough has nothing to walk through. Go back and capture the key tasks first.

Step 2: Dispatch the heuristic-evaluator Agent

Dispatch the heuristic-evaluator agent (see agents/heuristic-evaluator.md). It will:

  1. Evaluate the interface against Nielsen's 10 heuristics, citing specific violations with evidence (H1–H10)
  2. Run a cognitive walkthrough of every key task — four questions per step
  3. Analyse error paths, not just happy paths — recovery, undo, back navigation, dead ends
  4. Assess learnability (first-time user) and efficiency (repeat user)
  5. Acknowledge what works well, not only what fails

The agent narrates at three moments (arrival, working, departure) per the Agent Transparency protocol in using-designpowers.

Step 3: Run in Parallel and Reconcile

Per the Reconciliation Protocol in using-designpowers, the heuristic-evaluator runs simultaneously with design-critic and the accessibility-reviewer:

design-builder finishes
        |
   ┌────┼────────┐
   v    v        v
critic  reviewer  heuristic    (run simultaneously)
   |    |        |
   └────┼────────┘
        v
  reconciliation

When findings overlap or conflict, classify them (Aligned / Complementary / Conflicting) and resolve using the protocol's priority rules. Remember: usability wins over style — a beautiful interface that confuses people has failed.

Show full SKILL.md (225 more words)Show less
Step 4: Honour the Auto-Mode Safeguard

Even in auto mode, the pipeline must pause if the heuristic-evaluator finds a critical H3 violation (no undo on a destructive action) or H1 violation (user is completely lost). These indicate structural problems, not polish issues, and the user should decide how to resolve them.

⚠️ Auto paused: heuristic-evaluator found a critical usability issue (H3 — no undo on delete) that needs your decision. [details]

Step 5: Feed the Fix Round

After reconciliation:

  1. The agent's findings join the single prioritised fix list (critical first)
  2. Note which findings the heuristic-evaluator sourced and whether any were reconciled with other reviewers
  3. Record deferred Minor findings in the Design Debt Register via design-debt-tracker — promises don't disappear because severity is low
  4. Re-run the heuristic-evaluator only on critical fixes after the build round, not the full evaluation

What You Get

A structured heuristic evaluation report (see the agent definition for the full format): a heuristic findings table (H1–H10 with verdicts), cognitive walkthrough results per task, findings grouped by severity, what works well, and a recommendation (Proceed / Revise / Rethink).

Integration

  • Runs after: design-builder (build complete) — or against an existing design in Review Mode
  • Runs alongside: designpowers-critique, accessibility review (parallel, then reconciled)
  • Feeds: the reconciliation protocol → design-builder fix round → synthetic-user-testing
  • Calls: design-debt-tracker for deferred Minor findings
  • Records to: design-state.md (Decisions Log, Handoff Chain, Open Questions)

© Owl-Listener, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/heuristic-evaluation of Owl-Listener/designpowers.

Open the folder on GitHubat commit cb00757

Compare with similar skills

Heuristic Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Heuristic Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Heuristic Evaluation this skillOwl-Listener/designpowers251—~1.3kAutomated safety check: PassMIT
Color Auditrome-os/rome737—~2.7kAutomated safety check: PassMIT
Dont Make Me Thinkluongnv89/skills131—~2.5kAutomated safety check: PassMIT
Super UI UX Designsyahiidkamil/Software-Engineer-AI-Agent-Atlas401—~9kAutomated safety check: PassNone
Design Critiquegetcrew44/crew44356—~691Automated safety check: PassMIT
Error Prevention and RecoveryOwl-Listener/inclusive-design-skills104—~739Automated safety check: PassMIT

Similar skills

  • Color Audit

    rome-os/rome

    Audit a design system's color palette against measurable color-science disciplines — WCAG/APCA contrast of declared token pairs, perceptual (OKLCH) ramp uniformity, color-blindness safety of…

    737 GitHub stars~2.7k tokensUpdated today
    Frontend & DesignAuto-check passed
  • Dont Make Me Think

    luongnv89/skills

    Review UI usability using Steve Krug's principles and produce a scannable report.

    131 GitHub stars~2.5k tokensUpdated today
    Frontend & DesignAuto-check passed
  • Super UI UX Design

    syahiidkamil/Software-Engineer-AI-Agent-Atlas

    Complete UI/UX design capability — a wiki-style knowledge map of design theory (UX laws, Nielsen heuristics, usability, visual hierarchy, typography, contrast/WCAG, design systems…

    401 GitHub stars~9k tokensUpdated 3 mo ago
    Frontend & DesignAuto-check passed
  • Design Critique

    getcrew44/crew44

    Walks an agent through five review passes on a screen, flow or mockup, then returns ranked findings with a severity and a suggested fix for each one.

    356 GitHub stars~691 tokensUpdated 4 mo ago
    Frontend & DesignAuto-check passed
  • Error Prevention and Recovery

    Owl-Listener/inclusive-design-skills

    Designs forgiving forms and flows: prevent input errors, write messages that say what happened and what to do, and add undo, confirmation and recovery paths.

    104 GitHub stars~739 tokensUpdated 4 mo ago
    Frontend & DesignAuto-check passed
  • UI Automation Workflows

    conorluddy/xclaude-plugin

    Accessibility-first UI automation using IDB. An agent skill from conorluddy/xclaude-plugin.

    183 GitHub stars~2.2k tokensUpdated 27 days ago
    Frontend & DesignAuto-check passed

More from Owl-Listener/designpowers

All 33 skills in this repo
  • Adaptive Interfaces

    Owl-Listener/designpowers

    A skill your agent uses when designing for user preferences — motion sensitivity, contrast needs, colour schemes, text sizing, information density, or any interface behaviour that should adapt to…

    251 GitHub stars~1.1k tokensUpdated 3 mo ago
    Auto-check passed
  • Design Debate

    Owl-Listener/designpowers

    A skill your agent uses when a design direction is uncertain, when the team could go multiple ways, or when the user wants to see competing approaches argued before committing — orchestrates…

    251 GitHub stars~2.1k tokensUpdated 3 mo ago
    Auto-check passed
  • Design Debt Tracker

    Owl-Listener/designpowers

    A skill your agent uses when critique or review produces deferred findings, when checking accumulated design compromises, or when deciding what to address in the next iteration.

    251 GitHub stars~2.2k tokensUpdated 3 mo ago
    Auto-check passed
  • Design Discovery

    Owl-Listener/designpowers

    You MUST use this before any creative or design work — building features, creating components, designing interfaces, modifying user-facing behaviour.

    251 GitHub stars~2.1k tokensUpdated 3 mo ago
    Auto-check passed
  • Design Handoff

    Owl-Listener/designpowers

    A skill your agent uses when design work is complete and needs to be communicated to engineering — creates specifications, documents rationale, accessibility requirements, and interaction details in…

    251 GitHub stars~1.1k tokensUpdated 3 mo ago
    Auto-check passed
  • Design Review

    Owl-Listener/designpowers

    A skill your agent uses when the user wants to evaluate something that ALREADY EXISTS rather than build something new — "review this", "audit this screen", "what's wrong with this page", "is this…

    251 GitHub stars~1.7k tokensUpdated 3 mo ago
    Auto-check passed

Questions about Heuristic Evaluation

What does Heuristic Evaluation do?

Use after a build to evaluate usability against Nielsen's 10 heuristics and run cognitive walkthroughs of every key task. Heuristic Evaluation is an agent skill from Owl-Listener/designpowers. Use after a build to evaluate usability against Nielsen's 10 heuristics and run cognitive walkthroughs of every key task.

When should I use Heuristic Evaluation?

Heuristic Evaluation fits situations like: tasks that involve UX design; tasks that involve Accessibility.

How do I install Heuristic Evaluation in Claude Code?

Run `npx skills add Owl-Listener/designpowers --skill heuristic-evaluation -a claude-code`. Or copy the skill folder (skills/heuristic-evaluation in Owl-Listener/designpowers) into .claude/skills/heuristic-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Heuristic Evaluation in Codex?

Run `npx skills add Owl-Listener/designpowers --skill heuristic-evaluation -a codex`. Or copy the skill folder (skills/heuristic-evaluation in Owl-Listener/designpowers) into .agents/skills/heuristic-evaluation in your project. Codex loads it when a task matches its description.

Can I use Heuristic Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Owl-Listener/designpowers --skill heuristic-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/heuristic-evaluation, .gemini/skills/heuristic-evaluation, .github/skills/heuristic-evaluation and .opencode/skills/heuristic-evaluation in your project.

What does Heuristic Evaluation need to run?

SKILL.md names no scripts, command-line tools or credentials: Heuristic Evaluation is instructions for the agent only.

Does Heuristic Evaluation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Heuristic Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Heuristic Evaluation use?

Heuristic Evaluation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Heuristic Evaluation use?

About 1.3k tokens (SKILL.md is roughly 5.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Heuristic Evaluation?

Skills that share tags, products or a category with Heuristic Evaluation: Color Audit (rome-os/rome, 737 stars), Dont Make Me Think (luongnv89/skills, 131 stars), Super UI UX Design (syahiidkamil/Software-Engineer-AI-Agent-Atlas, 401 stars) and Design Critique (getcrew44/crew44, 356 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Heuristic Evaluation?

Owl-Listener (a GitHub user) maintains it in Owl-Listener/designpowers, which has 251 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on June 23, 2026.

Source: Owl-Listener/designpowers on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.