Agent skill

Ulw QA

by rlaope in rlaope/oh-my-hermes

[omh] Hostile scenario testing: adversarial QA and fix loops.

MITAuto-check passedTesting & QA

Install Ulw QA

skills CLI
$ npx skills add rlaope/oh-my-hermes --skill ulw-qa -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install rlaope/oh-my-hermes ulw-qa --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/rlaope/oh-my-hermes.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ulw-qa .claude/skills/ulw-qa && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ulw-qa
GitHub stars
3.2k
Token cost
~2.2k tokens
SKILL.md length
1,161 words
Files
3 (incl. references)
Skills in repo
143
Repo updated
First seen
Licence
MIT

At a glance

[omh] Hostile scenario testing: adversarial QA and fix loops.

  • The user says: ultraqa
  • SKILL.md covers Why This Exists, Do Not Use When, Examples and Completion Checklist, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Hostile scenarios

What it does

Ulw QA is an agent skill from rlaope/oh-my-hermes. [omh] Hostile scenario testing: adversarial QA and fix loops. Use when the user says: ultraqa, adversarial qa, hostile scenarios, e2e qa, real-world qa, qa scenario, release qa, 敵対的QA.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/board-fanin.md` and `references/manual-test-guide.md`).

It sits in Testing & QA, covering End-to-end testing. The repository describes itself as: All in one plugin for Hermes Agent ⚚ the coding intelligence, a long-term memory system and model optimized workflow packages. The licence is MIT.

When your agent uses it

  • The user says: ultraqa
  • Hostile scenarios

Example prompts

  • “/ulw-qa”

What it can do on your machine

Read from SKILL.md and the folder at commit 41de9dc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ulw QA loads about 2.2k tokens when it runs, and up to ~5.7k if it reads all its reference files. Until then it costs about 48 tokens; SKILL.md has 1,161 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~48
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from rlaope/oh-my-hermes at commit 41de9dc, republished under its MIT licence (© rlaope). 1,161 words, ~2,176 tokens.

Download SKILL.mdSave it as .claude/skills/ulw-qa/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
ulw-qa
description
[omh] Hostile scenario testing: adversarial QA and fix loops. Use when the user says: ultraqa, adversarial qa, hostile scenarios, e2e qa, real-world qa, qa scenario, release qa, 敵対的QA.

Ultraqa

This is a Hermes-native ultraqa workflow skill.

Why This Exists

ultraqa exists to keep verification work explicit, evidence-backed, and inside the Hermes/executor boundary instead of relying on ad hoc chat narration.

Do Not Use When

  • The request is casual chat, a status-only acknowledgement, or another workflow has stronger routing evidence.
  • The user needs implementation, review, CI, merge, or external publishing evidence that has not been delegated or observed.

Examples

Good example:

  • Prompt: $ultraqa test the setup wizard with hostile install paths, stale config, and missing PATH cases.
  • Expected behavior: Generate adversarial QA scenarios, expected signals, observed results, and fix-or-retry routing.
  • Why: The request asks for verification pressure and hostile scenarios.

Bad example:

  • Prompt: ultraqa: treat casual chat or unaccepted work as if this workflow already produced verified results.
  • Expected behavior: Ask a clarification question or route to a narrower workflow instead of forcing ultraqa.
  • Why: The request lacks the required inputs or would overclaim work that Hermes did not observe.

Completion Checklist

  • The scenario, expected behavior, observed result, and pass/fail basis are named.
  • Proposed fixes are separated from observed QA evidence.
  • Missing or failed verification routes back to plan, fix, or a narrower test.

Recovery Notes

  • If the expected behavior is unclear, route back to plan before running adversarial checks.
  • If verification fails, return to fix or research with the failed signal instead of advancing.

Workflow Lane

  • Current lane: Coding handoff (idea-to-deploy, llm-app-dev, cto-loop, deploy-and-monitor, code-review, build-failure-triage, verification-gate, security-safety-review, +28 more) - coding owners, handoffs, review, CI, and merge evidence.
  • If intent belongs to another lane, hand back to oh-my-hermes or name the adjacent workflow.
  • Shared product, routing, compatibility, and evidence rules: omh-routing/references/skill-common-rail.md.

Use When

Use when the task needs adversarial test scenarios, verification, and fix loops.

Strong routing signals: `ultraqa`, `$ultraqa`, `adversarial qa`, `hostile scenarios`, `e2e qa`, `real-world qa`, `qa scenario`, `release qa`, `敵対的QA`, `リリース前QA`, `障害シナリオ`, `장애 상황`, `쿠버네티스 장애`, `적절히 진단`, `검증 체크리스트`, `릴리즈 전 gate`, `对抗式测试`, `发布前测试`, `故障场景`

Catalog Metadata

Category: verification Phase: qa Hermes role: reviewer Quality tier: scenario-gated Reasoning demand: standard

Quality bar:

  • Do not start this engine as an automatic continuation of another skill's output: an accepted plan, a clarified brief, or a routing recommendation is planning evidence, not permission. Unless the user explicitly invoked this engine themselves, restate in one line what will start (engine, scope, selected executor) and wait for the user's explicit go-ahead first.
  • A mid-run user message is an interjection, not a stop: answer it briefly and, in the same reply, continue the run — re-read the phase todo when one is active and dispatch or advance the next pending step, or name the armed wait it is waiting on -- handle, bound completion signal, deadline -- instead of re-reading status. Only the user's explicit stop or cancel, or the engine's own completion gate, ends the run; when the interjection changes scope, say so and update the declared plan or todo instead of silently abandoning it. A mid-run message is the latest steering for the active task, not automatically a replacement objective: it replaces the objective when the user says so and steers the current one otherwise.
  • A follow-up that needs new authority, materially expands the scope, or changes external state not already authorized is described first and started only on the user's approval: the turn ends by naming that next action and asking whether to take it, as one question carrying the choices the user has, never by declaring what will not be done; persistence never broadens the authorized scope. A refused escalation is answered the same way, with a safer alternative inside the boundary or the authorization the boundary asks for — never a workaround or an indirect execution.
  • The closing brief scales to the change: one or two sentences plus the observed validation for a simple change, more only when the complexity earns it. Lead with the result or decision, in the user's words; omit abandoned approaches unless they explain a tradeoff the reader needs; narrate no internal bookkeeping (todo transitions, waits). When the work stops at a boundary or at a decision the user owns, end with the next action offered as a question, and state what was left undone as the option it leaves open, never as a refusal. Required closing lines stay outside this scaling: the observed run summary, and any prepared-not-observed or unmerged work, are stated whatever the brief's length.
  • Generate hostile scenarios from changed behavior and known risk areas.
  • Report pass/fail evidence separately from proposed fixes.
  • For native omh_todo checkpoints, load the todo-checklist closing recipe; record then recall this qa declaration. Stored declarations are not proof.
  • Delegate code mutations discovered by QA to the selected coding executor.
  • A check no automation can hold gets a manual test guide, never a done claim: load references/manual-test-guide.md for the step shape (setup, do, expect, broken), the four reasons a step may stay manual, and why it stays prepared_not_observed until a run records a fresh observed_check_results/v1. code-story phase VI. Manual test guide; rendered surfaces go to visual-qa.
  • For probes that must run in Hermes-owned isolation or outlive this session, load references/board-fanin.md: one probe row per scenario in its own worktree, one fixer row whose parents is every probe, re-verification through the review lane, and findings only through bounded readback.
  • When Hermes owns the coding path, read hermes_coding_harness/v1 before saying build, verification, review, docs, or PR-prep evidence exists.
Show full SKILL.md (274 more words)Show less

Handoff policy:

Hermes can design scenarios and report observed results; code fixes discovered by QA should become selected executor/runtime handoffs.

Required inputs:

  • changed behavior
  • acceptance criteria
  • known risk areas

Expected outputs:

  • adversarial scenarios
  • pass/fail evidence
  • fix recommendations

Artifact expectations:

  • QA scenario evidence
  • runtime verification summary

Safety rules:

  • Do not imply hidden Hermes runtime behavior.
  • Use the smallest verification that can prove the claim.

Runtime Evidence

Preferred harness for this skill: qa-specialist.

sh
omh runtime record --skill ultraqa --harness qa-specialist --status started

Record observed delegation results; otherwise return not_available or not_observed. Prepared OMH routing is not execution, review, CI, merge-readiness, or merge evidence.

  • When wrapper metadata includes memory_review_card/v1 or handoff_context_pack/v1, treat it as reviewed OMH-local or wrapper-supplied context only. Use conflict-free context summaries to shape plans and handoffs, but do not claim Hermes internal memory was read or changed. Preserve workflow intent and stop conditions; verify before claiming completion. Reply in the user's own words and the host's own voice: its SOUL.md persona owns reply language, tone, speech level, and sentence endings, progress updates included (where it sets no language, use the one the user wrote in), and OMH shapes structure and content only; OMH's record terms (surface, lane, wrapper, handoff, evidence boundary, not_observed) stay in records and tool calls, never in the sentence the user reads unless they ask about one; and when a stop condition or a decision the user owns ends the turn, offer the next action as a question rather than declaring what will not be done.

Use Hermes-native subagent/delegation features when available: native subagents -> Hermes delegation when available, otherwise sequential lanes.

Shared product, compatibility, topology, memory, harness, and execution rules: omh-routing/references/skill-common-rail.md. Load it when applicable; otherwise name an unavailable capability.

© rlaope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/ulw-qa of rlaope/oh-my-hermes.

  • SKILL.md
  • references/board-fanin.md
  • references/manual-test-guide.md

Open the folder on GitHubat commit 41de9dc

Compare with similar skills

Ulw QA next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ulw QA compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ulw QA this skillrlaope/oh-my-hermes3.2k—~2.2kAutomated safety check: PassMIT
Web Application Testinganthropics/skills180k51 repos~966Automated safety check: PassApache-2.0
TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph11211 repos~2.4kAutomated safety check: PassNone
Uloop Replay Inputkurotu/VRCQuestTools3733 repos~615Automated safety check: PassMIT
Ui4 Convert Testspayloadcms/payload45k—~3.5kAutomated safety check: PassMIT
E2Estackia/rtp2httpd2.2k—~517Automated safety check: PassGPL-2.0

Similar skills

  • Web Application Testing

    anthropics/skills

    Official

    Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

    180k GitHub starsUsed in 51 repos~966 tokens
    Testing & QAAuto-check passed
  • TDD Workflow

    hellangleZ/burn-in-cceverywhere-ralph

    A skill your agent uses when writing new features, fixing bugs, or refactoring code.

    112 GitHub starsUsed in 11 repos~2.4k tokens
    Testing & QAAuto-check passed
  • Uloop Replay Input

    kurotu/VRCQuestTools

    Replay recorded PlayMode keyboard and mouse input. An agent skill from kurotu/VRCQuestTools.

    373 GitHub starsUsed in 3 repos~615 tokens
    Testing & QAAuto-check passed
  • Ui4 Convert Tests

    payloadcms/payload

    A skill your agent uses when UI changes are complete and e2e tests need updating.

    45k GitHub stars~3.5k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • E2E

    stackia/rtp2httpd

    Write, run, review, or debug rtp2httpd E2E tests and their harness in e2e/ and scripts/run-e2e.sh.

    2.2k GitHub stars~517 tokensUpdated 7 days ago
    Testing & QAAuto-check passed
  • Moav E2E

    MotherofallVPNs/MoaV

    Run and debug MoaV's end-to-end tests — real protocol connectivity (client-test.sh) and the moav CLI smoke test — against a LIVE server, via the self-hosted e2e workflow or a local test VPS.

    448 GitHub stars~1.9k tokensUpdated 2 days ago
    Testing & QAAuto-check: notes

More from rlaope/oh-my-hermes

All 143 skills in this repo
  • Omh Accessibility Audit

    rlaope/oh-my-hermes

    [omh] Screen-reader or keyboard accessibility gaps: prepare WCAG, keyboard, focus, screen-reader, target-size, and reflow evidence gates for UI surfaces.

    3.2k GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed
  • Omh Agent Evaluation

    rlaope/oh-my-hermes

    [omh] Choosing between coding agents on evidence: compare executor or agent choices on reproducible tasks using quality, cost, time, tool, and evidence metrics.

    3.2k GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Omh Agent Instructions

    rlaope/oh-my-hermes

    [omh] Agent instruction file for a repo -- AGENTS.md, CLAUDE.md, a Cursor rule: write or update what an agent cannot derive from the code, inside a marked region, with every command verified or…

    3.2k GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Omh Agent Ops Review

    rlaope/oh-my-hermes

    [omh] AI agent progress for managers: help managers inspect AI-agent progress, blockers, quality gates, and throughput levers.

    3.2k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Omh AI Slop Cleaner

    rlaope/oh-my-hermes

    [omh] Messy or AI-generated code to clean up: delete AI-generated slop, dead code, and duplication while observable behavior stays identical.

    3.2k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check passed
  • Omh App Debugging

    rlaope/oh-my-hermes

    [omh] Application code misbehaves -- a wrong value, a flaky test, a lost update: reproduce it first, form competing hypotheses, discriminate them with the cheapest observation, and only then fix the…

    3.2k GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Ulw QA

What does Ulw QA do?

[omh] Hostile scenario testing: adversarial QA and fix loops. Ulw QA is an agent skill from rlaope/oh-my-hermes. [omh] Hostile scenario testing: adversarial QA and fix loops.

When should I use Ulw QA?

Ulw QA fits situations like: the user says: ultraqa; hostile scenarios.

How do I install Ulw QA in Claude Code?

Run `npx skills add rlaope/oh-my-hermes --skill ulw-qa -a claude-code`. Or copy the skill folder (skills/ulw-qa in rlaope/oh-my-hermes) into .claude/skills/ulw-qa in your project. Claude Code loads it when a task matches its description.

How do I install Ulw QA in Codex?

Run `npx skills add rlaope/oh-my-hermes --skill ulw-qa -a codex`. Or copy the skill folder (skills/ulw-qa in rlaope/oh-my-hermes) into .agents/skills/ulw-qa in your project. Codex loads it when a task matches its description.

Can I use Ulw QA in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rlaope/oh-my-hermes --skill ulw-qa -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ulw-qa, .gemini/skills/ulw-qa, .github/skills/ulw-qa and .opencode/skills/ulw-qa in your project.

What does Ulw QA need to run?

SKILL.md names no scripts, command-line tools or credentials: Ulw QA is instructions for the agent only.

Does Ulw QA access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ulw QA safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ulw QA use?

Ulw QA is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ulw QA use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.6k tokens, read only when the agent opens those files.

What are the alternatives to Ulw QA?

Skills that share tags, products or a category with Ulw QA: Web Application Testing (anthropics/skills, 180k stars), TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars), Uloop Replay Input (kurotu/VRCQuestTools, 373 stars) and Ui4 Convert Tests (payloadcms/payload, 45k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ulw QA?

rlaope (a GitHub user) maintains it in rlaope/oh-my-hermes, which has 3,233 GitHub stars. The repository holds 143 skills in this directory. The repository was last updated on October 8, 2026.

Source: rlaope/oh-my-hermes on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.