Agent skill

Synthetic User Testing

by Owl-Listener in Owl-Listener/designpowers

Use after the fix round to validate the design by walking through key tasks as each persona — simulating how Jordan (low-vision), Priya (non-native speaker), Marcus (motor impairment), or any…

MITAuto-check passedProduct & Project Management

Install Synthetic User Testing

skills CLI
$ npx skills add Owl-Listener/designpowers --skill synthetic-user-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Owl-Listener/designpowers synthetic-user-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Owl-Listener/designpowers.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/synthetic-user-testing .claude/skills/synthetic-user-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
synthetic-user-testing
GitHub stars
251
Token cost
~2.3k tokens
SKILL.md length
819 words
Files
1
Skills in repo
33
Repo updated
First seen
Licence
MIT

At a glance

Use after the fix round to validate the design by walking through key tasks as each persona — simulating how Jordan (low-vision), Priya (non-native speaker), Marcus (motor impairment), or any…

  • Works in 6 steps: Gather Test Context → Define Test Scenarios → Walk Through as Each Persona → …
  • Tasks that involve User research
  • SKILL.md covers When to Use, Process, What You Deliver and Integration, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Synthetic User Testing is an agent skill from Owl-Listener/designpowers. Use after the fix round to validate the design by walking through key tasks as each persona — simulating how Jordan (low-vision), Priya (non-native speaker), Marcus (motor impairment), or any project persona would actually experience the interface. Catches the issues that code review misses because they only surface in the act of using

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Product & Project Management, covering User research. The repository describes itself as: An agent design team you control: 10 agents that run an inclusive design process while you direct. The licence is MIT.

When your agent uses it

  • Tasks that involve User research

Example prompts

  • “/synthetic-user-testing”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Gather Test Context
  2. Define Test Scenarios
  3. Walk Through as Each Persona
  4. Cross-Persona Analysis
  5. Synthesise Findings
  6. Feed Results Forward

What it can do on your machine

Read from SKILL.md and the folder at commit cb00757. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Synthetic User Testing loads about 2.3k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 819 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Owl-Listener/designpowers at commit cb00757, republished under its MIT licence (© Owl-Listener). 819 words, ~2,281 tokens.

Download SKILL.mdSave it as .claude/skills/synthetic-user-testing/SKILL.md (or your agent's skills folder).
name
synthetic-user-testing
description
Use after the fix round to validate the design by walking through key tasks as each persona — simulating how Jordan (low-vision), Priya (non-native speaker), Marcus (motor impairment), or any project persona would actually experience the interface. Catches the issues that code review misses because they only surface in the act of using

Synthetic User Testing

Synthetic user testing is the closest thing to putting the design in front of real people without leaving the pipeline. You walk through the interface as each persona, attempting real tasks, and report what works, what breaks, and what feels wrong — from their perspective, not yours.

This is not a checklist exercise. It is an act of empathy disciplined by specifics. When you test as Jordan (low-vision, uses 200% zoom and high contrast), you don't ask "is contrast sufficient?" — you ask "can Jordan find the 'continue reading' button at 200% zoom when three articles are competing for attention?"

When to Use

  • After the fix round — the design-builder has addressed findings from critic, accessibility-reviewer, and heuristic-evaluator. Now validate that the fixes work and no new issues were introduced
  • Before verification-before-shipping — synthetic testing feeds directly into the persona walkthrough in the verification report
  • When the design-critic flags persona coverage gaps — if the critique says "unclear whether Persona X can complete Task Y," run a synthetic test to find out
  • When the team is unsure about a flow — synthetic testing turns "I think this works for screen reader users" into "here's exactly where a screen reader user would get stuck"

Process

Step 1: Gather Test Context

Before testing, assemble:

  • Personas from inclusive-personas (via design-state.md)
  • Key tasks from the design brief — the things the design must enable
  • The build — test the actual implementation, not the spec
  • Assistive technology context — for each persona, what tools they use and how (zoom level, screen reader, keyboard-only, switch access, etc.)
Step 2: Define Test Scenarios

For each key task in the brief, write a scenario that a persona would actually encounter. Scenarios are not "test case 1" — they are moments:

Format:

SCENARIO: [Natural situation trigger]
TASK: [What the persona is trying to accomplish]
PERSONA: [Name] — [key ability context]
SUCCESS: [What "done" looks like for this persona]

Example:

SCENARIO: Jordan is on the train home and wants to finish an article
          they started yesterday.
TASK: Find and continue a partially-read article.
PERSONA: Jordan — low vision, uses 200% zoom, high contrast mode,
         reads on a phone with one hand.
SUCCESS: Jordan locates the article, picks up where they left off,
         and the progress updates when they finish.
Step 3: Walk Through as Each Persona

For each scenario, simulate the persona's experience step by step. This is not "would this work?" — it is "let me try to do this the way [Persona] would."

For each step, document:

STEP [N]: [What the persona does]
USING: [Input method — touch, keyboard, screen reader, switch, etc.]
SEES: [What the interface presents — at their zoom level, contrast
       setting, screen size]
THINKS: [What the persona would likely think or feel]
RESULT: ✓ succeeds / ⚠ succeeds with difficulty / ✗ fails / ? unclear

FINDING: [If ⚠, ✗, or ? — what went wrong and why]
WHO IS AFFECTED: [This persona, and any others with similar needs]

Critical rules for walking through:

  1. Stay in character — if the persona uses a screen reader, evaluate what the screen reader would announce, not what the sighted experience looks like
  2. Use their device — if the persona reads on a phone at 200% zoom, evaluate at that zoom on a phone viewport
  3. Include emotional state — a stressed parent and a relaxed commuter approach the same interface differently. The scenario sets the emotional context
  4. Test error paths — what happens if the persona makes a mistake? Can they recover? How much does recovery cost them?
  5. Note friction, not just failure — a task that technically succeeds but requires 8 taps and 3 scrolls has a friction problem even if it "works"
Step 4: Cross-Persona Analysis

After walking through all scenarios, look for patterns across personas:

Barrier matrix:

| Task              | Jordan    | Priya     | Marcus    | [Persona] |
|                   | (low vis) | (ESL)     | (motor)   |           |
|-------------------|-----------|-----------|-----------|-----------|
| Find article      | ⚠ zoom    | ✓         | ✗ target  | ...       |
| Continue reading  | ✓         | ⚠ jargon  | ✓         | ...       |
| Mark complete     | ✗ no fbk  | ✓         | ⚠ gesture | ...       |

Pattern analysis:

  • Universal barriers — issues that affect 2+ personas → likely a design problem, not an edge case
  • Persona-specific barriers — issues that affect only one persona → may need targeted fix or alternative path
  • Friction hotspots — steps where multiple personas struggle, even if they eventually succeed
  • Emotional patterns — where do personas feel confused, frustrated, or lost?
Show full SKILL.md (302 more words)Show less
Step 5: Synthesise Findings

Compile findings into a structured report (see "What You Deliver" below). Every finding must:

  • Name the specific persona affected
  • Describe the exact step where the issue occurs
  • Explain why it's a problem from the persona's perspective
  • Suggest a fix
  • Classify severity (Critical / Major / Minor)

Severity in synthetic testing:

SeverityDefinition
CriticalPersona cannot complete the task at all
MajorPersona completes the task but with significant difficulty, confusion, or emotional friction
MinorPersona completes the task but the experience is rougher than it should be
Step 6: Feed Results Forward

Synthetic testing results feed into two places:

  1. design-builder — if fixes are needed, dispatch the builder with specific findings
  2. verification-before-shipping — the persona walkthrough section of the verification report should reference synthetic test results, not guesswork

What You Deliver

markdown
# Synthetic User Test Results: [Project Name]

**Date:** [YYYY-MM-DD]
**Build tested:** [what was tested]
**Personas tested:** [list]
**Tasks tested:** [list]

## Summary
[2-3 sentences: overall findings — who can use this, who can't,
 and the biggest gap]

## Scenario Results

### Scenario 1: [Scenario name]
**Persona:** [Name] — [context]
**Task:** [what they're trying to do]
**Result:** ✓ / ⚠ / ✗

[Step-by-step walkthrough with findings]

### Scenario 2: ...

## Barrier Matrix

| Task | [Persona 1] | [Persona 2] | [Persona 3] | ... |
|------|-------------|-------------|-------------|-----|
| ...  | ✓/⚠/✗      | ✓/⚠/✗      | ✓/⚠/✗      | ... |

## Cross-Persona Patterns
- **Universal barriers:** [issues affecting 2+ personas]
- **Friction hotspots:** [steps with clustered difficulty]
- **Emotional patterns:** [where confusion/frustration clusters]

## Findings by Severity

### Critical
- [Persona] cannot [task] because [specific reason] → [Fix]

### Major
- [Persona] struggles with [task] at [step] because [reason] → [Fix]

### Minor
- [Persona] experiences friction at [step] because [reason] → [Fix]

## Comparison: Pre-Fix vs Post-Fix
[If this is a re-test after fixes, show what improved and what didn't]

## Recommendation
[Ship / Fix and re-test / Rethink flow for [persona]]

Integration

  • Runs after: Fix round (design-builder has addressed critic, accessibility-reviewer, and heuristic-evaluator findings)
  • Runs before: verification-before-shipping
  • Informed by: inclusive-personas, design-discovery (brief), design-state (all decisions)
  • Feeds into: verification-before-shipping (persona walkthrough evidence), design-builder (if fixes needed), design-debt-tracker (deferred findings)
  • Complements: usability-testing (which plans tests with real people — synthetic testing validates before real testing begins)

The Difference from Usability Testing

Synthetic User TestingUsability Testing
WhoAI walks through as each personaReal people use the interface
WhenAfter fix round, inside the pipelineAfter shipping or during user research
SpeedMinutesDays to weeks
CatchesPredictable barriers, flow breaks, frictionUnexpected behaviour, mental model mismatches, emotional reactions
MissesTrue surprises — things no persona model predictsNothing (but expensive and slow)
ValueCloses the gap between build and validation without leaving the pipelineGround truth

Synthetic testing does not replace real usability testing. It makes real testing more efficient by catching the obvious issues first, so real participants encounter the design at its best — and surface the surprises only humans can find.

© Owl-Listener, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/synthetic-user-testing of Owl-Listener/designpowers.

Open the folder on GitHubat commit cb00757

Compare with similar skills

Synthetic User Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Synthetic User Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Synthetic User Testing this skillOwl-Listener/designpowers251—~2.3kAutomated safety check: PassMIT
Fable DomainSahir619/fable-method2.3k—~2.6kAutomated safety check: PassMIT
Design Sprintwondelai/skills2.4k—~3.8kAutomated safety check: PassMIT
User Research Cookiycookiy-ai/user-research-skill1.6k—~954Automated safety check: PassMIT
Produck Feedback To Buildtryproduck/produck-skills510—~1kAutomated safety check: PassApache-2.0
Customer InterviewsRefoundAI/lenny-skills1.4k—~1.7kAutomated safety check: PassMIT

Similar skills

  • Fable Domain

    Sahir619/fable-method

    Discuss a domain with the user, research it from real sources, then generate a trusted skill bundle for it - a step-by-step workflow with a flowchart, a domain adapter, a trap fixture, and a smoke…

    2.3k GitHub stars~2.6k tokensUpdated 6 days ago
    Product & Project ManagementAuto-check passed
  • Design Sprint

    wondelai/skills

    Run a structured 5-day process to prototype, test, and validate product ideas with real users.

    2.4k GitHub stars~3.8k tokensUpdated 28 days ago
    Product & Project ManagementAuto-check passed
  • User Research Cookiy

    cookiy-ai/user-research-skill

    End-to-end user research assistant — qualitative and quantitative.

    1.6k GitHub stars~954 tokensUpdated 1 mo ago
    Product & Project ManagementAuto-check passed
  • Produck Feedback To Build

    tryproduck/produck-skills

    Pulls full in-context user feedback tickets through the Produck MCP server and turns them into an aligned product change instead of a guess.

    510 GitHub stars~1k tokensUpdated 1 mo ago
    Product & Project ManagementAuto-check passed
  • Customer Interviews

    RefoundAI/lenny-skills

    Help users conduct high-impact customer interviews that move beyond surface-level feature requests to identify root emotional frustrations and specific causal triggers.

    1.4k GitHub stars~1.7k tokensUpdated 2 mo ago
    Product & Project ManagementAuto-check passed
  • Guides a product discovery conversation and writes product-brief.md with the problem, evidence, scope, decisions and the next open question, for existing, client or own ideas.

    231 GitHub stars~3k tokensUpdated 4 days ago
    Product & Project ManagementAuto-check passed

More from Owl-Listener/designpowers

All 33 skills in this repo
  • Adaptive Interfaces

    Owl-Listener/designpowers

    A skill your agent uses when designing for user preferences — motion sensitivity, contrast needs, colour schemes, text sizing, information density, or any interface behaviour that should adapt to…

    251 GitHub stars~1.1k tokensUpdated 3 mo ago
    Auto-check passed
  • Design Debate

    Owl-Listener/designpowers

    A skill your agent uses when a design direction is uncertain, when the team could go multiple ways, or when the user wants to see competing approaches argued before committing — orchestrates…

    251 GitHub stars~2.1k tokensUpdated 3 mo ago
    Auto-check passed
  • Design Debt Tracker

    Owl-Listener/designpowers

    A skill your agent uses when critique or review produces deferred findings, when checking accumulated design compromises, or when deciding what to address in the next iteration.

    251 GitHub stars~2.2k tokensUpdated 3 mo ago
    Auto-check passed
  • Design Discovery

    Owl-Listener/designpowers

    You MUST use this before any creative or design work — building features, creating components, designing interfaces, modifying user-facing behaviour.

    251 GitHub stars~2.1k tokensUpdated 3 mo ago
    Auto-check passed
  • Design Handoff

    Owl-Listener/designpowers

    A skill your agent uses when design work is complete and needs to be communicated to engineering — creates specifications, documents rationale, accessibility requirements, and interaction details in…

    251 GitHub stars~1.1k tokensUpdated 3 mo ago
    Auto-check passed
  • Design Review

    Owl-Listener/designpowers

    A skill your agent uses when the user wants to evaluate something that ALREADY EXISTS rather than build something new — "review this", "audit this screen", "what's wrong with this page", "is this…

    251 GitHub stars~1.7k tokensUpdated 3 mo ago
    Auto-check passed

Questions about Synthetic User Testing

What does Synthetic User Testing do?

Use after the fix round to validate the design by walking through key tasks as each persona — simulating how Jordan (low-vision), Priya (non-native speaker), Marcus (motor impairment), or any…. Synthetic User Testing is an agent skill from Owl-Listener/designpowers. Use after the fix round to validate the design by walking through key tasks as each persona — simulating how Jordan (low-vision), Priya (non-native speaker), Marcus (motor impairment), or any project persona would actually experience the interface.

When should I use Synthetic User Testing?

Synthetic User Testing fits situations like: tasks that involve User research.

How do I install Synthetic User Testing in Claude Code?

Run `npx skills add Owl-Listener/designpowers --skill synthetic-user-testing -a claude-code`. Or copy the skill folder (skills/synthetic-user-testing in Owl-Listener/designpowers) into .claude/skills/synthetic-user-testing in your project. Claude Code loads it when a task matches its description.

How do I install Synthetic User Testing in Codex?

Run `npx skills add Owl-Listener/designpowers --skill synthetic-user-testing -a codex`. Or copy the skill folder (skills/synthetic-user-testing in Owl-Listener/designpowers) into .agents/skills/synthetic-user-testing in your project. Codex loads it when a task matches its description.

Can I use Synthetic User Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Owl-Listener/designpowers --skill synthetic-user-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/synthetic-user-testing, .gemini/skills/synthetic-user-testing, .github/skills/synthetic-user-testing and .opencode/skills/synthetic-user-testing in your project.

What does Synthetic User Testing need to run?

SKILL.md names no scripts, command-line tools or credentials: Synthetic User Testing is instructions for the agent only.

Does Synthetic User Testing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Synthetic User Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Synthetic User Testing use?

Synthetic User Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Synthetic User Testing use?

About 2.3k tokens (SKILL.md is roughly 9.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Synthetic User Testing?

Skills that share tags, products or a category with Synthetic User Testing: Fable Domain (Sahir619/fable-method, 2.3k stars), Design Sprint (wondelai/skills, 2.4k stars), User Research Cookiy (cookiy-ai/user-research-skill, 1.6k stars) and Produck Feedback To Build (tryproduck/produck-skills, 510 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Synthetic User Testing?

Owl-Listener (a GitHub user) maintains it in Owl-Listener/designpowers, which has 251 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on June 23, 2026.

Source: Owl-Listener/designpowers on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.