Agent skill

Usability Testing

by rampstackco in rampstackco/claude-skills

Plan and run usability tests on existing or prototype designs including test design, task scripts, moderation, observation, and findings synthesis.

MITAuto-check passedProduct & Project Management

Install Usability Testing

skills CLI
$ npx skills add rampstackco/claude-skills --skill usability-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install rampstackco/claude-skills usability-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/rampstackco/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/usability-testing .claude/skills/usability-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
usability-testing
GitHub stars
935
Token cost
~2.4k tokens
SKILL.md length
1,112 words
Files
3 (incl. references)
Skills in repo
103
Repo updated
First seen
Licence
MIT

At a glance

Plan and run usability tests on existing or prototype designs including test design, task scripts, moderation, observation, and findings synthesis.

  • Works in 5 steps: Define what to test → Choose moderated or unmoderated → Recruit → …
  • The user wants to test usability
  • SKILL.md covers When to use, When NOT to use, Required inputs and The framework: 5 phases, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Usability Testing is an agent skill from rampstackco/claude-skills. Plan and run usability tests on existing or prototype designs including test design, task scripts, moderation, observation, and findings synthesis. Use this skill whenever the user wants to test usability, run a moderated test, run an unmoderated test, validate a design, find usability issues, or improve task completion. Triggers on usability test, usability testing, moderated test, unmoderated test, task script, think aloud, prototype testing, user testing, design validation, task completion. Also triggers when…

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `README.md` and `references/task-script-patterns.md`).

It sits in Product & Project Management, covering User research and UX design. The repository describes itself as: Stack-agnostic Claude Skills covering the full website lifecycle: brand, design, content, SEO, dev, ops, growth, and research. Build, ship, audit, optimize. The licence is MIT.

When your agent uses it

  • The user wants to test usability
  • Run a moderated test
  • Run an unmoderated test
  • Validate a design

Example prompts

  • “/usability-testing”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Define what to test
  2. Choose moderated or unmoderated
  3. Recruit
  4. Run the test
  5. Synthesize and report

What it can do on your machine

Read from SKILL.md and the folder at commit 482c9bf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Usability Testing loads about 2.4k tokens when it runs, and up to ~5.1k if it reads all its reference files. Until then it costs about 156 tokens; SKILL.md has 1,112 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~156
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from rampstackco/claude-skills at commit 482c9bf, republished under its MIT licence (© rampstackco). 1,112 words, ~2,389 tokens.

Download SKILL.mdSave it as .claude/skills/usability-testing/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
usability-testing
description
Plan and run usability tests on existing or prototype designs including test design, task scripts, moderation, observation, and findings synthesis. Use this skill whenever the user wants to test usability, run a moderated test, run an unmoderated test, validate a design, find usability issues, or improve task completion. Triggers on usability test, usability testing, moderated test, unmoderated test, task script, think aloud, prototype testing, user testing, design validation, task completion. Also triggers when the user has built something and wants to know if real users can use it before shipping.
category
research
catalog_summary
Test design, moderation, findings reports
display_order
2

Usability Testing

Plan and run tests that find usability problems before users hit them in production. Stack-agnostic. Tool-agnostic.

This skill is for testing existing designs or prototypes. For broader discovery research, use ux-research. For conversion testing in production, use cro-optimization.


When to use

  • Before launching a new flow or major redesign
  • After a redesign to verify it doesn't introduce new problems
  • When analytics show drop-off but you don't know why
  • When customer support tickets pattern around specific UI areas
  • Pre-launch user validation
  • Comparing two design directions

When NOT to use

  • Discovery / generative research (use ux-research)
  • Live conversion optimization (use cro-optimization)
  • Mapping the broader experience (use journey-mapping)
  • Pure quantitative measurement (use analytics-strategy)

Required inputs

  • The design or prototype to test (functional or near-functional)
  • Specific tasks users would do
  • The audience (who should be tested)
  • Testing infrastructure (moderated tool, unmoderated tool, in-person setup)

The framework: 5 phases

1. Define what to test

Don't test the whole product. Test specific tasks.

Task selection criteria:

  • The task represents a real user goal (not "click around and explore")
  • The task has a clear start and end
  • The task is achievable in 2 to 10 minutes
  • The task is one of: most common, most strategic, most problematic

Examples of testable tasks:

"You want to find a contractor near you who can install a fence. Show me how you'd do that on this site."

"You're a first-time visitor. You want to understand if this product fits your needs. Walk me through how you'd evaluate it."

"Your team needs a new tool to manage projects. Use this site to figure out which plan is right for a 12-person team."

Task framing rules:

  • State the user goal, not the system action ("find a place to stay" not "click the search button")
  • Provide context (why are you doing this?)
  • Don't reveal the path
  • Don't use product terminology in the task framing
2. Choose moderated or unmoderated

Moderated (live, with researcher):

  • Researcher observes and probes in real time
  • Best for early-stage prototypes, complex tasks, novel concepts
  • Higher cost, smaller sample (5 to 8 participants typical)
  • Catches surprises and probe deeper

Unmoderated (recorded, asynchronous):

  • Participant completes alone, often via tool (UserTesting, Maze, Lookback)
  • Best for stable designs, simple tasks, larger sample
  • Lower cost, larger sample (15 to 30 participants typical)
  • Catches patterns at scale, less depth per session

For most teams: moderated for early/critical decisions, unmoderated for ongoing validation.

3. Recruit

Target audience - not just convenience.

Recruit criteria:

  • Match real users (target audience, not just "anyone")
  • Mix of experience levels with the product (new and existing if applicable)
  • Mix of relevant device types (mobile, desktop, tablet if relevant)
  • Exclude friends, family, employees

Sample size:

  • Moderated: 5 to 8 participants (Nielsen's "5 users find 85% of usability issues" for the most common segment)
  • Unmoderated: 15 to 30 participants (more participants compensate for less probing)
  • Multi-segment testing: 5 to 8 per segment
4. Run the test

Pre-task setup:

  • Confirm recording works
  • Brief participant (purpose, anonymity, recording, "no wrong answers")
  • Get verbal consent
  • Have participant share screen if remote

Moderated session structure:

  1. Warm-up (2 to 3 min). Easy questions to put participant at ease.
  2. Pre-test questions (3 to 5 min). Background context, current behavior with similar products.
  3. Task 1 (5 to 10 min). Describe task. Have participant attempt while thinking aloud.
  4. Post-task questions (1 to 2 min). What was easy/hard? Anything confusing?
  5. Repeat for tasks 2, 3, 4 (typically 3 to 5 tasks per 60-minute session).
  6. Overall debrief (5 to 10 min). General reactions, comparisons to alternatives, anything else.
  7. Close (2 min).

Moderation principles:

  • Encourage think-aloud ("What's going through your mind?")
  • Don't help unless they're truly stuck (and even then, only after a long pause)
  • Don't lead ("Are you looking for the menu?" - bad)
  • Note where they hesitate, scroll, or backtrack
  • Note their language vs the product's language
  • Note emotional reactions

Anti-patterns:

  • Talking too much (researcher should talk maybe 20% of the time)
  • Defending the design when participants struggle
  • Helping prematurely
  • Asking participants to predict their future behavior
  • Treating participant suggestions as features ("Users want X" - test demand for X separately)
Show full SKILL.md (429 more words)Show less
5. Synthesize and report

Patterns across participants are signal. Single-participant complaints are weaker (but worth investigating).

Synthesis steps:

  1. Issue inventory. Every issue observed, with which participant, which task, severity.
  2. Cluster. Issues that are the same root problem.
  3. Severity.
    • Critical: Blocks task completion. Most users hit this.
    • Major: Significantly slows task. Many users hit this.
    • Minor: Friction. Some users hit this. Workaround exists.
    • Cosmetic: Polish. Doesn't affect task.
  4. Recommendations. For each issue, propose specific fixes.
  5. Prioritize. By severity and effort.

Report structure:

markdown
# Usability Test: [Design / flow]

## Summary
[2 to 3 paragraphs covering: what was tested, headline findings, top 3 priorities]

## Method
[Moderated/unmoderated, sample size, audience, dates, tasks]

## Critical findings
[Each with description, frequency, supporting evidence (quotes/clips), recommendation. Where evidence was not obtained, state the gap per the data-availability rule]

## Major findings
[Same structure]

## Minor findings
[Brief]

## Cosmetic findings
[Briefest]

## What worked well
[Calibration: capture successes too]

## Recommendations
[Prioritized list with effort estimates]

## Next steps
[Test re-run schedule, design iteration plan]

Workflow

  1. Define the goals. What decisions hinge on this? What tasks matter most?
  2. Design tasks. 3 to 5 specific, realistic, goal-framed tasks.
  3. Choose moderated vs unmoderated. Match to stage and depth needed.
  4. Recruit. Specific to audience.
  5. Pilot. 1 to 2 sessions before main batch. Refine tasks if needed.
  6. Run. Follow the protocol. Stay disciplined.
  7. Synthesize during, not just after. Patterns emerge by session 4 or 5.
  8. Report. Multiple formats - written report + highlight clips.
  9. Track fixes. Every critical issue should have an owner and date.
  10. Re-test after fixes. Verify the fix worked, didn't introduce new issues.

Failure patterns

  • Testing the whole product instead of specific tasks. Vague results.
  • Tasks that reveal the path. ("Click the menu and find...")
  • Friends and family as participants. Biased, not representative.
  • Researcher leading the participant. Findings reflect the researcher.
  • Defending the design when participants struggle. Misses real issues.
  • Helping too quickly. Participant doesn't experience the friction.
  • Treating participant suggestions as features. Users solve their problem; product team designs the solution.
  • One participant = data point. A single strong opinion isn't a finding.
  • Skipping severity scoring. All findings treated equally; team can't prioritize.
  • Reports no one reads. Highlight clips and live walkthroughs work better than 80-page decks.
  • Testing once, never re-testing. Fixes that introduce new problems go undetected.

Output format

Default outputs:

  1. Test plan (before testing) - usability-test-plan-[topic].md
  2. Task script (per session) - usability-tasks-[topic].md
  3. Findings report (after synthesis) - usability-findings-[topic].md
  4. Highlight clips (separately produced)

If required data is unavailable

This skill's output depends on data, measurements, or tool results it cannot generate on its own. When a required input, tool, or data source is unavailable or unverifiable, the sanctioned output is the deliverable with the gap stated: what was needed, what was actually obtained or verified, and which parts of the output are affected. Fabricating, estimating, or interpolating a required number to complete the deliverable is never sanctioned. A stated gap is a complete answer.


Reference files

© rampstackco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/usability-testing of rampstackco/claude-skills.

  • SKILL.md
  • README.md
  • references/task-script-patterns.md

Open the folder on GitHubat commit 482c9bf

Compare with similar skills

Usability Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Usability Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Usability Testing this skillrampstackco/claude-skills935—~2.4kAutomated safety check: PassMIT
Lean UX Canvas v2deanpeters/Product-Manager-Skills7.2k1 repos~6.2kAutomated safety check: PassCustom licence
UX Researcher Designeralirezarezvani/claude-skills28k1 repos~3kAutomated safety check: PassMIT
Product Researchalirezarezvani/claude-skills28k—~2.7kAutomated safety check: PassMIT
Building ProductGTM-Strategist/gtm-strategist-skills261—~5.8kAutomated safety check: PassMIT
Investigateghaida/intent206—~7.3kAutomated safety check: PassCC0-1.0

Similar skills

  • Lean UX Canvas v2

    deanpeters/Product-Manager-Skills

    Guides a team through Jeff Gothelf's Lean UX Canvas v2 to frame a business problem, surface assumptions and decide what to learn and test next.

    7.2k GitHub starsUsed in 1 repo~6.2k tokens
    Product & Project ManagementAuto-check passed
  • UX Researcher Designer

    alirezarezvani/claude-skills

    UX research and design toolkit for Senior UX Designer/Researcher including data-driven persona generation, journey mapping, usability testing frameworks, and research synthesis.

    28k GitHub starsUsed in 1 repo~3k tokens
    Product & Project ManagementAuto-check passed
  • Product Research

    alirezarezvani/claude-skills

    A skill your agent uses when planning and synthesizing product/user research as a method-and-repository discipline — selecting the right method for the goal (generative interviews vs usability test…

    28k GitHub stars~2.7k tokensUpdated 1 mo ago
    Product & Project ManagementAuto-check passed
  • Building Product

    GTM-Strategist/gtm-strategist-skills

    A skill your agent uses when the user needs to define their MVP, create a product roadmap, plan metrics and tracking, refine their value proposition with JTBD, or run usability tests.

    261 GitHub stars~5.8k tokensUpdated 2 mo ago
    Product & Project ManagementAuto-check passed
  • Investigate

    ghaida/intent

    Guide and conduct user research — from planning through synthesis.

    206 GitHub stars~7.3k tokensUpdated 2 mo ago
    Product & Project ManagementAuto-check passed
  • Discover Interview Synthesis

    product-on-purpose/pm-skills

    Synthesizes user research interviews into actionable insights, patterns, and recommendations.

    713 GitHub stars~1.3k tokensUpdated 3 days ago
    Product & Project ManagementAuto-check passed

More from rampstackco/claude-skills

All 103 skills in this repo
  • After Action Report

    rampstackco/claude-skills

    Run a structured after-action review (postmortem, retrospective) on a launch, incident, or completed project to capture timeline, root cause analysis, contributing factors, and actionable lessons.

    935 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Analytics Strategy

    rampstackco/claude-skills

    Design measurement frameworks including event taxonomy, KPI hierarchy, dashboard architecture, attribution models, and analytics implementation strategy.

    935 GitHub starsUsed in 1 repo~2.4k tokens
    Auto-check passed
  • Brand Style Guide

    rampstackco/claude-skills

    Build or audit a comprehensive brand style guide that documents the full brand system including story, logo system, color, typography, imagery, voice, applications, and dos/don'ts.

    935 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Brand Voice

    rampstackco/claude-skills

    Develop or document a complete brand voice and tone system covering voice attributes, tone shifts by context, vocabulary preferences, grammar rules, and copy examples.

    935 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Content And Copy

    rampstackco/claude-skills

    Write or edit website copy, blog content, and editorial pieces with attention to voice, structure, and goal.

    935 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Content Strategy

    rampstackco/claude-skills

    Develop a content strategy covering editorial positioning, content pillars, formats, calendar, governance, and topical authority planning.

    935 GitHub stars~2.6k tokensUpdated today
    Auto-check passed

Questions about Usability Testing

What does Usability Testing do?

Plan and run usability tests on existing or prototype designs including test design, task scripts, moderation, observation, and findings synthesis. Usability Testing is an agent skill from rampstackco/claude-skills. Plan and run usability tests on existing or prototype designs including test design, task scripts, moderation, observation, and findings synthesis.

When should I use Usability Testing?

Usability Testing fits situations like: the user wants to test usability; run a moderated test; run an unmoderated test; validate a design.

How do I install Usability Testing in Claude Code?

Run `npx skills add rampstackco/claude-skills --skill usability-testing -a claude-code`. Or copy the skill folder (skills/usability-testing in rampstackco/claude-skills) into .claude/skills/usability-testing in your project. Claude Code loads it when a task matches its description.

How do I install Usability Testing in Codex?

Run `npx skills add rampstackco/claude-skills --skill usability-testing -a codex`. Or copy the skill folder (skills/usability-testing in rampstackco/claude-skills) into .agents/skills/usability-testing in your project. Codex loads it when a task matches its description.

Can I use Usability Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rampstackco/claude-skills --skill usability-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/usability-testing, .gemini/skills/usability-testing, .github/skills/usability-testing and .opencode/skills/usability-testing in your project.

What does Usability Testing need to run?

SKILL.md names no scripts, command-line tools or credentials: Usability Testing is instructions for the agent only.

Does Usability Testing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Usability Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Usability Testing use?

Usability Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Usability Testing use?

About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.7k tokens, read only when the agent opens those files.

What are the alternatives to Usability Testing?

Skills that share tags, products or a category with Usability Testing: Lean UX Canvas v2 (deanpeters/Product-Manager-Skills, 7.2k stars), UX Researcher Designer (alirezarezvani/claude-skills, 28k stars), Product Research (alirezarezvani/claude-skills, 28k stars) and Building Product (GTM-Strategist/gtm-strategist-skills, 261 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Usability Testing?

rampstackco (a GitHub organization) maintains it in rampstackco/claude-skills, which has 935 GitHub stars. The repository holds 103 skills in this directory. The repository was last updated on October 7, 2026.

Source: rampstackco/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.