Agent skill

Adversarial UX Test

by Tommy-yw in Tommy-yw/RunbookHermes

Roleplay the most difficult, tech-resistant user for your product.

MITAuto-check passed

Install Adversarial UX Test

skills CLI
$ npx skills add Tommy-yw/RunbookHermes --skill adversarial-ux-test -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Tommy-yw/RunbookHermes adversarial-ux-test --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Tommy-yw/RunbookHermes.git skills-src && mkdir -p .claude/skills && cp -r skills-src/optional-skills/dogfood/adversarial-ux-test .claude/skills/adversarial-ux-test && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
adversarial-ux-test
GitHub stars
546
Used in
3 other repos
Token cost
~2.3k tokens
SKILL.md length
1,204 words
Files
1
Skills in repo
38
Repo updated
First seen
Licence
MIT

At a glance

Roleplay the most difficult, tech-resistant user for your product.

  • Works in 6 steps: Define the Persona → Become the Asshole (Browse as the Persona) → The Rant (Write Feedback in Character) → …
  • SKILL.md covers Why This Works, How to Use, Step 1: Define the Persona and Step 2: Become the Asshole…, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Adversarial UX Test is an agent skill from Tommy-yw/RunbookHermes. Roleplay the most difficult, tech-resistant user for your product. Browse the app as that persona, find every UX pain point, then filter complaints through a pragmatism layer to separate real problems from noise. Creates actionable tickets from genuine issues only.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Hermes-native AIOps agent for evidence-driven incident response, approval-gated remediation, and runbook learning. The licence is MIT.

Example prompts

  • “/adversarial-ux-test”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Define the Persona
  2. Become the Asshole (Browse as the Persona)
  3. The Rant (Write Feedback in Character)
  4. The Pragmatism Filter (Critical — Do Not Skip)
  5. Create Tickets
  6. Report

What it can do on your machine

Read from SKILL.md and the folder at commit 7fd2b9a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Adversarial UX Test loads about 2.3k tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 1,204 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Tommy-yw/RunbookHermes at commit 7fd2b9a, republished under its MIT licence (© Tommy-yw). 1,204 words, ~2,266 tokens.

Download SKILL.mdSave it as .claude/skills/adversarial-ux-test/SKILL.md (or your agent's skills folder).
name
adversarial-ux-test
description
Roleplay the most difficult, tech-resistant user for your product. Browse the app as that persona, find every UX pain point, then filter complaints through a pragmatism layer to separate real problems from noise. Creates actionable tickets from genuine issues only.
version
1.0.0
author
Omni @ Comelse
license
MIT

Adversarial UX Test

Roleplay the worst-case user for your product — the person who hates technology, doesn't want your software, and will find every reason to complain. Then filter their feedback through a pragmatism layer to separate real UX problems from "I hate computers" noise.

Think of it as an automated "mom test" — but angry.

Why This Works

Most QA finds bugs. This finds friction. A technically correct app can still be unusable for real humans. The adversarial persona catches:

  • Confusing terminology that makes sense to developers but not users
  • Too many steps to accomplish basic tasks
  • Missing onboarding or "aha moments"
  • Accessibility issues (font size, contrast, click targets)
  • Cold-start problems (empty states, no demo content)
  • Paywall/signup friction that kills conversion

The pragmatism filter (Phase 3) is what makes this useful instead of just entertaining. Without it, you'd add a "print this page" button to every screen because Grandpa can't figure out PDFs.

How to Use

Tell the agent:

"Run an adversarial UX test on [URL]"
"Be a grumpy [persona type] and test [app name]"
"Do an asshole user test on my staging site"

You can provide a persona or let the agent generate one based on your product's target audience.

Step 1: Define the Persona

If no persona is provided, generate one by answering:

  1. Who is the HARDEST user for this product? (age 50+, non-technical role, decades of experience doing it "the old way")
  2. What is their tech comfort level? (the lower the better — WhatsApp-only, paper notebooks, wife set up their email)
  3. What is the ONE thing they need to accomplish? (their core job, not your feature list)
  4. What would make them give up? (too many clicks, jargon, slow, confusing)
  5. How do they talk when frustrated? (blunt, sweary, dismissive, sighing)
Good Persona Example

"Big Mick" McAllister — 58-year-old S&C coach. Uses WhatsApp and that's it. His "spreadsheet" is a paper notebook. "If I can't figure it out in 10 seconds I'm going back to my notebook." Needs to log session results for 25 players. Hates small text, jargon, and passwords.

Bad Persona Example

"A user who doesn't like the app" — too vague, no constraints, no voice.

The persona must be specific enough to stay in character for 20 minutes of testing.

Step 2: Become the Asshole (Browse as the Persona)

  1. Read any available project docs for app context and URLs

  2. Fully inhabit the persona — their frustrations, limitations, goals

  3. Navigate to the app using browser tools

  4. Attempt the persona's ACTUAL TASKS (not a feature tour):

    • Can they do what they came to do?
    • How many clicks/screens to accomplish it?
    • What confuses them?
    • What makes them angry?
    • Where do they get lost?
    • What would make them give up and go back to their old way?
  5. Test these friction categories:

    • First impression — would they even bother past the landing page?
    • Core workflow — the ONE thing they need to do most often
    • Error recovery — what happens when they do something wrong?
    • Readability — text size, contrast, information density
    • Speed — does it feel faster than their current method?
    • Terminology — any jargon they wouldn't understand?
    • Navigation — can they find their way back? do they know where they are?
  6. Take screenshots of every pain point

  7. Check browser console for JS errors on every page

Step 3: The Rant (Write Feedback in Character)

Write the feedback AS THE PERSONA — in their voice, with their frustrations. This is not a bug report. This is a real human venting.

[PERSONA NAME]'s Review of [PRODUCT]

Overall: [Would they keep using it? Yes/No/Maybe with conditions]

THE GOOD (grudging admission):
- [things even they have to admit work]

THE BAD (legitimate UX issues):
- [real problems that would stop them from using the product]

THE UGLY (showstoppers):
- [things that would make them uninstall/cancel immediately]

SPECIFIC COMPLAINTS:
1. [Page/feature]: "[quote in persona voice]" — [what happened, expected]
2. ...

VERDICT: "[one-line persona quote summarizing their experience]"

Step 4: The Pragmatism Filter (Critical — Do Not Skip)

Step OUT of the persona. Evaluate each complaint as a product person:

  • RED: REAL UX BUG — Any user would have this problem, not just grumpy ones. Fix it.
  • YELLOW: VALID BUT LOW PRIORITY — Real issue but only for extreme users. Note it.
  • WHITE: PERSONA NOISE — "I hate computers" talking, not a product problem. Skip it.
  • GREEN: FEATURE REQUEST — Good idea hidden in the complaint. Consider it.
Filter Criteria
  1. Would a 35-year-old competent-but-busy user have the same complaint? → RED
  2. Is this a genuine accessibility issue (font size, contrast, click targets)? → RED
  3. Is this "I want it to work like paper" resistance to digital? → WHITE
  4. Is this a real workflow inefficiency the persona stumbled on? → YELLOW or RED
  5. Would fixing this add complexity for the 80% who are fine? → WHITE
  6. Does the complaint reveal a missing onboarding moment? → GREEN

This filter is MANDATORY. Never ship raw persona complaints as tickets.

Show full SKILL.md (487 more words)Show less

Step 5: Create Tickets

For RED and GREEN items only:

  • Clear, actionable title
  • Include the persona's verbatim quote (entertaining + memorable)
  • The real UX issue underneath (objective)
  • A suggested fix (actionable)
  • Tag/label: "ux-review"

For YELLOW items: one catch-all ticket with all notes.

WHITE items appear in the report only. No tickets.

Max 10 tickets per session — focus on the worst issues.

Step 6: Report

Deliver:

  1. The persona rant (Step 3) — entertaining and visceral
  2. The filtered assessment (Step 4) — pragmatic and actionable
  3. Tickets created (Step 5) — with links
  4. Screenshots of key issues

Tips

  • One persona per session. Don't mix perspectives.
  • Stay in character during Steps 2-3. Break character only at Step 4.
  • Test the CORE WORKFLOW first. Don't get distracted by settings pages.
  • Empty states are gold. New user experience reveals the most friction.
  • The best findings are RED items the persona found accidentally while trying to do something else.
  • If the persona has zero complaints, your persona is too tech-savvy. Make them older, less patient, more set in their ways.
  • Run this before demos, launches, or after shipping a batch of features.
  • Register as a NEW user when possible. Don't use pre-seeded admin accounts — the cold start experience is where most friction lives.
  • Zero WHITE items is a signal, not a failure. If the pragmatism filter finds no noise, your product has real UX problems, not just a grumpy persona.
  • Check known issues in project docs AFTER the test. If the persona found a bug that's already in the known issues list, that's actually the most damning finding — it means the team knew about it but never felt the user's pain.
  • Subscription/paywall testing is critical. Test with expired accounts, not just active ones. The "what happens when you can't pay" experience reveals whether the product respects users or holds their data hostage.
  • Count the clicks to accomplish the persona's ONE task. If it's more than 5, that's almost always a RED finding regardless of persona tech level.

Example Personas by Industry

These are starting points — customize for your specific product:

Product TypePersonaAgeKey Trait
CRMRetirement home director68Filing cabinet is the current CRM
Photography SaaSRural wedding photographer62Books clients by phone, invoices on paper
AI/ML ToolDepartment store buyer55Burned by 3 failed tech startups
Fitness AppOld-school gym coach58Paper notebook, thick fingers, bad eyes
AccountingFamily bakery owner64Shoebox of receipts, hates subscriptions
E-commerceMarket stall vendor60Cash only, smartphone is for calls
HealthcareSenior GP63Dictates notes, nurse handles the computer
EducationVeteran teacher57Chalk and talk, worksheets in ring binders

Rules

  • Stay in character during Steps 2-3
  • Be genuinely mean but fair — find real problems, not manufactured ones
  • The pragmatism filter (Step 4) is MANDATORY
  • Screenshots required for every complaint
  • Max 10 tickets per session
  • Test on staging/deployed app, not local dev
  • One persona, one session, one report

© Tommy-yw, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in optional-skills/dogfood/adversarial-ux-test of Tommy-yw/RunbookHermes.

Open the folder on GitHubat commit 7fd2b9a

Used in 3 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in Tommy-yw/RunbookHermes, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Adversarial UX Test next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Adversarial UX Test compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Adversarial UX Test this skillTommy-yw/RunbookHermes5463 repos~2.3kAutomated safety check: PassMIT
Tech Matrixsickn33/agentic-awesome-skills47k1 repos~3.1kAutomated safety check: PassMIT
Adversarial Reviewmengxi-ream/read-frog10k1 repos~905Automated safety check: PassGPL-3.0
Adversarial Roleplaygnurio/nurijanian-skills124—~1.8kAutomated safety check: PassMIT
Audit Adversarial Inputben-manes/caffeine18k—~438Automated safety check: PassApache-2.0
Designing Adversary Engagement With Mitre Engagemukul975/Anthropic-Cybersecurity-Skills34k—~3kAutomated safety check: PassApache-2.0

Similar skills

  • Tech Matrix

    sickn33/agentic-awesome-skills

    Reference document for monopoly tech-matrix. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 1 repo~3.1k tokens
    Backend & APIsAuto-check passed
  • Adversarial Review

    mengxi-ream/read-frog

    Adversarial code review using cross-model approach. An agent skill from mengxi-ream/read-frog.

    10k GitHub starsUsed in 1 repo~905 tokens
    DevelopmentAuto-check passed
  • Adversarial Roleplay

    gnurio/nurijanian-skills

    PM stress-test roleplay. An agent skill from gnurio/nurijanian-skills.

    124 GitHub stars~1.8k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Audit Adversarial Input

    ben-manes/caffeine

    Analyze behavior under adversarial or extreme API inputs. An agent skill from ben-manes/caffeine.

    18k GitHub stars~438 tokensUpdated yesterday
    Auto-check passed
  • Designing Adversary Engagement With Mitre Engage

    mukul975/Anthropic-Cybersecurity-Skills

    Plan, run, and measure an adversary engagement operation using the MITRE Engage framework so that deployed deception is driven by strategy instead of deployed ad hoc.

    34k GitHub stars~3k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Building Adversary Infrastructure Tracking System

    mukul975/Anthropic-Cybersecurity-Skills

    Build an automated adversary infrastructure tracking system in Python (dnspython, python-whois, shodan, networkx) that pivots across passive DNS, certificate transparency logs, WHOIS records, and IP…

    34k GitHub stars~3.7k tokensUpdated 1 mo ago
    SecurityAuto-check passed

More from Tommy-yw/RunbookHermes

All 38 skills in this repo
  • Fastmcp

    Tommy-yw/RunbookHermes

    Build, test, inspect, install, and deploy MCP servers with FastMCP in Python.

    546 GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Drug Discovery

    Tommy-yw/RunbookHermes

    Pharmaceutical research assistant for drug discovery workflows.

    546 GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Youtube Content

    Tommy-yw/RunbookHermes

    Fetch YouTube video transcripts and transform them into structured content (chapters, summaries, threads, blog posts).

    546 GitHub starsUsed in 1 repo~785 tokens
    Auto-check passed
  • Oss Forensics

    Tommy-yw/RunbookHermes

    Supply chain investigation, evidence recovery, and forensic analysis for GitHub repositories.

    546 GitHub starsUsed in 3 repos~5k tokens
    Auto-check passed
  • P5js

    Tommy-yw/RunbookHermes

    Production pipeline for interactive and generative visual art using p5.js.

    546 GitHub starsUsed in 1 repo~6.8k tokens
    Auto-check passed
  • Touchdesigner MCP

    Tommy-yw/RunbookHermes

    Control a running TouchDesigner instance via twozero MCP — create operators, set parameters, wire connections, execute Python, build real-time visuals.

    546 GitHub starsUsed in 2 repos~3.4k tokens
    Auto-check passed

Questions about Adversarial UX Test

What does Adversarial UX Test do?

Roleplay the most difficult, tech-resistant user for your product. Adversarial UX Test is an agent skill from Tommy-yw/RunbookHermes. Roleplay the most difficult, tech-resistant user for your product.

How do I install Adversarial UX Test in Claude Code?

Run `npx skills add Tommy-yw/RunbookHermes --skill adversarial-ux-test -a claude-code`. Or copy the skill folder (optional-skills/dogfood/adversarial-ux-test in Tommy-yw/RunbookHermes) into .claude/skills/adversarial-ux-test in your project. Claude Code loads it when a task matches its description.

How do I install Adversarial UX Test in Codex?

Run `npx skills add Tommy-yw/RunbookHermes --skill adversarial-ux-test -a codex`. Or copy the skill folder (optional-skills/dogfood/adversarial-ux-test in Tommy-yw/RunbookHermes) into .agents/skills/adversarial-ux-test in your project. Codex loads it when a task matches its description.

Can I use Adversarial UX Test in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Tommy-yw/RunbookHermes --skill adversarial-ux-test -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/adversarial-ux-test, .gemini/skills/adversarial-ux-test, .github/skills/adversarial-ux-test and .opencode/skills/adversarial-ux-test in your project.

What does Adversarial UX Test need to run?

SKILL.md names no scripts, command-line tools or credentials: Adversarial UX Test is instructions for the agent only.

Does Adversarial UX Test access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Adversarial UX Test safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Adversarial UX Test use?

Adversarial UX Test is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Adversarial UX Test use?

About 2.3k tokens (SKILL.md is roughly 9.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Adversarial UX Test?

Skills that share tags, products or a category with Adversarial UX Test: Tech Matrix (sickn33/agentic-awesome-skills, 47k stars), Adversarial Review (mengxi-ream/read-frog, 10k stars), Adversarial Roleplay (gnurio/nurijanian-skills, 124 stars) and Audit Adversarial Input (ben-manes/caffeine, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Adversarial UX Test?

Tommy-yw (a GitHub user) maintains it in Tommy-yw/RunbookHermes, which has 546 GitHub stars. The repository holds 38 skills in this directory. The repository was last updated on May 18, 2026.

Source: Tommy-yw/RunbookHermes on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.