Agent skill

QA Strategy

by romiluz13 in romiluz13/cc10x

A skill your agent uses when designing or checking a QA test system: choosing the right tier, scenario matrices with pipeline-wide observation points, test-environment topology and isolation…

MITAuto-check: notesTesting & QA

Install QA Strategy

skills CLI
$ npx skills add romiluz13/cc10x --skill qa-strategy -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install romiluz13/cc10x qa-strategy --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/romiluz13/cc10x.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/cc10x/skills/qa-strategy .claude/skills/qa-strategy && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
qa-strategy
GitHub stars
164
Token cost
~5k tokens
SKILL.md length
2,923 words
Files
1
Skills in repo
22
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when designing or checking a QA test system: choosing the right tier, scenario matrices with pipeline-wide observation points, test-environment topology and isolation…

  • Works in 3 steps: enumerate the full input space (do this… → reduce with a technique, not by taste → chains, not just single actions
  • Checking a QA test system: choosing the right tier
  • SKILL.md covers Tier selection, The scenario matrix, Input-space coverage —… and Observation points — the…, plus 7 more sections
  • Calls docker

What it does

QA Strategy is an agent skill from romiluz13/cc10x. Use when designing or checking a QA test system: choosing the right tier, scenario matrices with pipeline-wide observation points, test-environment topology and isolation, fixture lifecycle, and flake sources. Also the coverage lens for reviewing a QA plan (it reaches a plan reviewer as a file to Read in the task scaffold, not as a SKILLHINTS entry).

Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Test strategy. The repository describes itself as: The Loop Engine for Claude Code — engineer the loop, not the prompt. 1 router · 9 agents · 16 skills · 4 workflows. Fail-closed gates, test honesty, anti-anchored review. The licence is MIT.

When your agent uses it

  • Checking a QA test system: choosing the right tier
  • Scenario matrices with pipeline-wide observation points
  • Test-environment topology and isolation
  • Fixture lifecycle

Example prompts

  • “/qa-strategy”

Requirements

  • Docker
  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Bash, LSP

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. enumerate the full input space (do this exhaustively)
  2. reduce with a technique, not by taste
  3. chains, not just single actions

What it can do on your machine

Read from SKILL.md and the folder at commit f346ebe. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Bash
    • LSP

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

QA Strategy loads about 5k tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 2,923 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~91
When it runs · the whole SKILL.md, loaded when a task matches
~5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Grep, Glob, Bash, LSP

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from romiluz13/cc10x at commit f346ebe, republished under its MIT licence (© romiluz13). 2,923 words, ~5,022 tokens.

Download SKILL.mdSave it as .claude/skills/qa-strategy/SKILL.md (or your agent's skills folder).
name
qa-strategy
description
Use when designing or checking a QA test system: choosing the right tier, scenario matrices with pipeline-wide observation points, test-environment topology and isolation, fixture lifecycle, and flake sources. Also the coverage lens for reviewing a QA plan (it reaches a plan reviewer as a file to Read in the task scaffold, not as a SKILL_HINTS entry).
allowed-tools
Read, Grep, Glob, Bash, LSP
user-invocable
false

QA Strategy

This skill is the QA route's test-system design discipline.

Core: A test suite's value is not how many tests it has. It is how much you would believe a green run.

Tier selection

The three-layer model in cc10x:building still governs — unit proves local behavior, integration proves boundary wiring, E2E proves system truth. QA owns the outer two, plus UI.

TierProvesUse when
integrationone boundary really worksroute→service, service→DB, producer→consumer, cache/side effects
e2e_backenda flow across services really worksthe value is produced by services cooperating
uia human can actually do the thingthe deliverable is something a person operates

Choose the cheapest tier that can still fail for the right reason. An E2E test that would also fail for ten unrelated reasons is a bad detector: it fires often and tells you little. Push a check down a tier whenever the lower tier can catch the same defect.

But do not push everything down. The defects that survive good unit coverage are exactly the ones that live between components — serialization mismatches, transaction boundaries, retry storms, partial failures. Those only appear at the tier where the components are real.

The scenario matrix

Every feature gets scenarios in three classes. A plan with only the first is not a test plan, it is a demo.

ClassQuestion it answers
happy-pathDoes the thing work when everything cooperates?
error-handlingWhen a dependency fails, does the system fail correctly — right status, right message, right rollback, right log?
edge-caseWhat about empty, one, many, max, concurrent, duplicate, out-of-order, expired, and unauthorized?
Error-handling scenarios are where the bugs are

For each boundary the flow crosses, ask: what happens when it is slow, down, returns garbage, returns a partial result, or succeeds after the caller gave up? Each answer that matters is a scenario.

The most valuable single question: when this fails halfway, what state is left behind? Partial-failure state is where data corruption lives, and it is almost never covered by accident.

Input-space coverage — enumerate exhaustively, then reduce deliberately

The goal is to cover as many user actions and chains of actions as possible: every option in a dropdown, every date class, every filter, every combination that a real user could produce.

Step 1 — enumerate the full input space (do this exhaustively)

For every interactive surface in the flow, list every option a user can actually pick:

  • every dropdown / select — every option, not "a representative one"
  • every filter, toggle, checkbox, radio
  • every date field — today, past, future, boundary of range, invalid, empty
  • every free-text field — its equivalence classes
  • every navigation path that reaches the same screen
  • every chain — the ordered sequences of actions a user strings together

Discover, do not assume. Options are usually data-driven; reading the component tells you a <select> exists, not what is in it at runtime. Drive the real UI to enumerate actual values (see UI tier tooling below). An enumeration built from the code alone will miss the options that only appear for certain roles, tenants, or feature flags.

Record the full enumeration in the test plan even when you will not execute all of it. The enumeration is the coverage claim. What you skip must be visible.

Step 2 — reduce with a technique, not by taste

Full combinatorial coverage is not achievable and pretending otherwise produces an unbuildable plan: 5 filters × 4 options each is 1,024 combinations; add a date and a sort and you are past 10,000. Nobody runs that suite, so in practice it never gets written, and you end up with less coverage than an honest reduction would have given.

Reduce with a named technique, and record which one you used and what it leaves uncovered:

TechniqueUse forWhat it costs
Every-option-oncedropdowns, radio groupsEach option exercised at least once. Misses interactions between options.
Pairwise (all-pairs)combinations of 3+ independent controlsEvery pair of values co-occurs in some case. Collapses 1,024 → ~25. Misses 3-way interactions, which are rare.
Equivalence classesfree text, numbers, datesOne representative per class of behavior. Misses defects inside a class.
Boundary valuesanything with a range0, 1, max, max+1, empty, null. Where most defects actually live.
Full combinatorial2 controls with few values, or a genuinely safety-critical pathComplete. Only affordable when the space is tiny.

Default recipe: every-option-once for single controls, pairwise across combinations, boundary values on every range, plus full combinatorial on any combination that is known to be risky (billing, permissions, anything where two settings interact by design).

Pairwise is the highest-leverage tool here. Most combination defects come from two settings interacting, not five. All-pairs buys the large majority of that value for a small fraction of the runs.

Step 3 — chains, not just single actions

A user action rarely stands alone. Cover the sequences too:

  • the primary chain — the full path a real user walks, end to end
  • back / re-entry — leave mid-flow and return; is state preserved or correctly discarded?
  • re-application — apply a filter, change it, clear it, apply another
  • order variation — does filter-then-search behave like search-then-filter?
  • repetition — the same action twice; the same submit double-clicked

Chains are where state bugs live, and they are almost never found by testing actions in isolation.

Feature flags are a coverage dimension, not a setup detail

A feature shipped behind a flag has two live code paths in production, and only one of them is new. Both need coverage:

StateWhat it provesWhy it gets skipped
Flag ONThe feature worksNobody skips this one
Flag OFFThe old path still works, or the feature is cleanly absentFeels redundant — it is not
Toggled mid-sessionState written under one path is readable under the otherRarely considered at all

The OFF path is where the expensive incidents live. A flag gets rolled back precisely when something is already going wrong, and that is the worst possible moment to discover the rollback does not restore a working system. If a flag is genuinely one-way, that is a finding to surface, not a row to skip.

Where entitlements or plan tiers exist alongside flags, they are independent gates. flag-on + entitlement-off must degrade correctly rather than 500 — a combination almost no plan covers, because each gate is usually tested alone.

Confirm a flag applied, do not confirm it was set. Asserting that the config call returned 200 proves the call returned 200. Caches, restart requirements, and per-pod config all break the leap from "set" to "in effect", and they break it silently.

Edge cases worth a default look
  • boundaries: 0, 1, max, max+1
  • concurrency: two of the same request at once; the same idempotency key twice
  • ordering: out-of-order and duplicate messages
  • time: expiry, timezone, clock skew, retry after timeout
  • identity: unauthorized, wrong tenant, expired token
  • data: unicode, very long strings, null vs. absent

Observation points — the pipeline-wide assertion

A scenario that asserts only the final response tests the response. It does not test the pipeline.

For each scenario, name what should be observable at every stage:

PointAsserts
UIwhat a person sees
APIstatus, response shape, headers
DBrows/documents actually written, in the right state
Queuemessages actually published, with the right payload
Logseach pipeline stage actually ran, and said so properly
Why logs are a first-class assertion

Logs are the only cheap way to prove that a middle stage executed. A 200 at the edge is consistent with a worker that never ran, a retry that silently swallowed, or a branch that fell through.

Asserting logs also makes the suite an enforcement point for the project's logging standard: if a scenario asserts a structured log line with named fields and the line is missing, unstructured, or logged at the wrong level, that is a real finding about observability — found by a test, before an incident needs it.

Assert on level + message + structured fields, not on a substring of a formatted line. Substring assertions on log text are among the most brittle tests it is possible to write.

Log access strategy. How the harness reads logs differs sharply by environment (local stdout, container logs, a log platform); state the access method in the plan for the environment in use.

Test environment

Isolation model

Ranked by strength:

  1. Ephemeral everything — fresh DB, fresh services per run. Strongest; slowest.
  2. Fresh data, shared services — services stay up, schema/data reset between runs. Usually the right trade.
  3. Namespaced in a shared environment — every run scopes itself by a unique prefix/tenant. Only when 1 and 2 are impossible; leaks are hard to detect and the failure mode is cross-run interference that looks like flake.

Never test against an environment someone else is using. A suite that intermittently fails because a colleague was clicking around teaches the team to ignore red.

Readiness, not sleeping

sleep 30 is a race condition with a comment. Gate on a real signal — health endpoint 200, port accepting, migration complete, topic created — with a bounded timeout and a loud failure.

An under-gated environment produces failures that look like product bugs. That is precisely how a team learns to distrust its own suite.

Teardown

Teardown must verify itself. "Ran docker compose down" is not evidence; "docker ps shows nothing from this run" is. A suite that leaks will eventually make its own machine unable to run it.

Determining the topology — ask, then detect, then confirm on screen

Never silently infer how to run the system. The order is fixed:

  1. Ask the user first. They know whether there is a compose file that works, a staging environment, or a cloud-only deploy. One question here saves an hour of wrong inference.
  2. Then detect, using whatever the machine actually offers — compose/Tilt/Skaffold files, testcontainers in the manifest, kind/k3d, available MCP servers, project skills, plugins, and agents that already know how to run this system.
  3. Then present the findings on screen and get explicit confirmation before building anything. Show the proposed topology in the conversation — not only written to a file. A plan the user never actually looked at is not an approved plan, and the environment is the single most expensive thing to get wrong.

Write the confirmed topology to env-plan.md after the user confirms it, not before.

Show full SKILL.md (1,230 more words)Show less
Settle contradictions with a probe, not with a preference

When the feature map records a contradiction about something the plan depends on — which tenant a record lands under, which queue is consumed, which identity is resolved, which of two documented shapes the wire actually carries — the environment can usually answer it, and a document cannot.

Write the first scenario as a probe: drive the smallest real path that exposes the disputed value, read it back, bind it to a plan variable, and enumerate one branch per outcome including a branch where neither expected answer appears and the run stops. Then place the probe in the bring-up sequence as a gate, before anything that consumes what it binds.

Picking the likelier answer instead is what makes contradictions expensive. The plan does not fail at the contradiction — it fails ten scenarios later as a mount fault, an empty queue, or a broker error, and the run gets spent debugging the harness instead of the product.

A stub inside the process under test can be the bug

Some controls can only be applied by injecting something into the process you are measuring — a clock shim, a module preload, a patched global, a stubbed transport. That is sometimes the only lever available, and it is legitimate. It is also a stub inside the thing under test, which means a failure it causes is indistinguishable from a product failure until someone looks.

When you must do it: name it as a risk in the plan, keep the patch as narrow as the assertion requires, and prefer patching a value over patching a mechanism. Patching Date.now moves thresholds. Patching the timer wheel moves the event loop into orderings the product never produces — the first is an observation aid, the second is a source of fiction.

And prove it applied. A shim that silently failed to load produces a scenario that passes for the wrong reason.

UI tier tooling

Two different tools for two different jobs. Use both, for what each is good at.

ToolUse it forDo not use it for
Agentic browser (claude-in-chrome)Discovery — driving the real UI to enumerate what is actually there: every dropdown option, every filter value, what a role can see, what a page really rendersThe deliverable. A browsing session is not an artifact; it cannot re-run next month without an agent and a live browser
PlaywrightThe deliverable — durable spec files that run headless, in CI, unattended, repeatablyExploring an unfamiliar UI from scratch; it is slow going without knowing what is on the page

The pipeline is: explore agentically → enumerate → codify as Playwright specs. Discovery feeds the input-space enumeration above; the specs are what qa-execute runs as a standalone regression suite forever after.

Selector durability matters more than it seems. Specs pinned to CSS classes or DOM position break on the next refactor and get deleted rather than fixed. Prefer roles, labels, and test ids — what a user perceives, not how it is currently marked up.

Fixtures and determinism

  • Deterministic by construction: fixed seeds, injected clock, controlled id generation. If a test can pass or fail depending on the wall clock, it will eventually do both.
  • Build data through the system's own front door where practical — a fixture written directly into the DB can encode a state the application can never actually produce, and then you are testing fiction.
  • Fixtures are code. They get reviewed. A fixture that drifts from the real schema produces tests that pass against a shape production never sends — the schema-incomplete-mock failure integration-verifier already watches for.
  • Where a fixture LIVES can change what the product does to it. Many systems exclude paths by convention — anything under test/, fixtures/, samples/, examples/, node_modules/, a .dockerignore entry, a gitignore-derived skip list. A fixture scanned, indexed, or uploaded from such a path is silently skipped, and the scenario under-reports with no error and no log line. Stage fixtures at a path with none of those components, and prove the product actually consumed the fixture rather than inferring it from a green result.

Flake sources, ranked

  1. Time — sleeps, timeouts tuned to one machine, timezone, clock skew
  2. Order dependence — a test that only passes after another test ran
  3. Shared state — leftover rows, caches, singletons, ports
  4. Real network — an external call in a test is a scheduled outage
  5. Under-gated readiness — asserting before the system is up

A flaky test is a broken test. Re-run once to classify it, never repeatedly to reach green. Quarantine and fix; a suite people re-run until it passes is a suite that proves nothing.

Proving the suite can fail

Before believing any harness: break the thing under test on purpose and confirm the test goes red. At minimum once per tier.

An assertion that has never been observed failing is unproven, and unproven assertions are how a suite drifts into decoration. This is the single highest-value check in the whole discipline, and it takes minutes.

The inverse failure — asserting a state the code cannot produce. Sometimes a scenario is unfalsifiable in the other direction: the state it checks for is unreachable, so the assertion can never fire at all. A status value the code never writes, an error branch with no caller, a field only a mock ever populates.

When you find one, assert the absence instead — "no row is written" is a real, checkable property, and it is usually the behaviour that actually matters. Reaching the unreachable state by writing a fixture past the application's own validators is testing fiction: it may still earn a place (the UI's handling of a shape it might one day receive is worth knowing), but label it as fiction in the plan so nobody reads its PASS as evidence about the product.

The tell is a scenario whose setup has to bypass the system's front door to exist.

Anti-patterns

Anti-patternWhy it is wrong
Asserting only the final responseProves the edge, not the pipeline
expect(x).toBeTruthy()Passes for almost everything
sleep as readinessRace condition with a comment
Catch-and-continue in setup/teardownTurns a broken environment into a green run
Skipped tests reporting as passedA lie with good manners
Re-running until greenConverts a real defect into a statistic
Fixtures written straight into the DBCan encode states the app cannot produce
Testing against a shared live environmentInterference reads as flake; team learns to ignore red
Editing the test to make the run passDestroys the only thing QA produces

Artifact templates

Every QA artifact has a shipped skeleton. Copy it and fill in place — do not improvise document structure, and do not delete a section that does not apply (mark it N/A with a reason).

ArtifactTemplate
test-plan.md${CLAUDE_PLUGIN_ROOT}/templates/qa-test-plan.template.md
env-plan.md${CLAUDE_PLUGIN_ROOT}/templates/qa-env-plan.template.md
feature-map.md${CLAUDE_PLUGIN_ROOT}/templates/qa-feature-map.template.md (router-owned, inline consolidation)
setup.md${CLAUDE_PLUGIN_ROOT}/templates/qa-setup.template.md — environment-scoped, not per-run. Lives at .cc10x/qa/env/{env_key}/setup.md, append-only, and records only MEASURED facts. It is the counterpart to env-plan.md: the plan predicts the environment from source, this records what the machine actually said
report.md${CLAUDE_PLUGIN_ROOT}/templates/qa-report.template.md (router-seeded at qa-execute; qa-executor rewrites it whole with Write)
harness manifest${CLAUDE_PLUGIN_ROOT}/templates/live-harness.template.json

Why deletion is forbidden. The sections most often dropped are the ones that record what the plan does not do — the coverage-reduction table, known gaps, teardown verification, re-runnability. Those are precisely the sections an optimistic plan omits. A missing section reads as "nothing to report"; an N/A with a reason reads as a claim someone can challenge.

© romiluz13, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/cc10x/skills/qa-strategy of romiluz13/cc10x.

Open the folder on GitHubat commit f346ebe

Compare with similar skills

QA Strategy next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

QA Strategy compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
QA Strategy this skillromiluz13/cc10x164—~5kAutomated safety check: NotesMIT
Testing OpenLogi UIAprilNEA/OpenLogi23k—~1.1kAutomated safety check: PassApache-2.0
Testing Hashqlhashintel/hash1.7k—~1.9kAutomated safety check: PassAGPL-3.0
Dynamo Unit TestingDynamoDS/Dynamo2k—~622Automated safety check: PassApache-2.0
Designing TestsCloudAI-X/opencode-workflow275—~2.9kAutomated safety check: PassMIT
Openprd Test StrategyDavidLam-oss/obsidian-wechat-converter332—~578Automated safety check: PassMIT

Similar skills

  • Testing OpenLogi UI

    AprilNEA/OpenLogi

    Verifies OpenLogi's native GPUI interface with focused tests, the component gallery and a mock agent, choosing the evidence that fits each change.

    23k GitHub stars~1.1k tokensUpdated 4 days ago
    Testing & QAAuto-check passed
  • Testing Hashql

    hashintel/hash

    HashQL testing strategies including compiletest (UI tests), unit tests, and snapshot tests.

    1.7k GitHub stars~1.9k tokensUpdated today
    Testing & QAAuto-check passed
  • Dynamo Unit Testing

    DynamoDS/Dynamo

    Write comprehensive NUnit tests for the Dynamo codebase following Dynamo testing patterns, conventions, and architectural constraints.

    2k GitHub stars~622 tokensUpdated today
    Testing & QAAuto-check passed
  • Designing Tests

    CloudAI-X/opencode-workflow

    Guides test strategy, TDD/BDD approaches, test coverage planning, and testing best practices.

    275 GitHub stars~2.9k tokensUpdated 9 mo ago
    Testing & QAAuto-check passed
  • Openprd Test Strategy

    DavidLam-oss/obsidian-wechat-converter

    OpenPrd 测试策略分流 skill:按风险把任务分到单元、集成、端到端、人工、视觉、小程序、性能和安全验证,并要求 evidence-plan。

    332 GitHub stars~578 tokensUpdated 5 days ago
    Testing & QAAuto-check passed
  • Go Testing

    Gentleman-Programming/gentle-ai

    Trigger: Go tests, go test coverage, Bubbletea teatest, golden files.

    7.6k GitHub stars~550 tokensUpdated today
    Testing & QAAuto-check passed

More from romiluz13/cc10x

All 22 skills in this repo
  • Building

    romiluz13/cc10x

    A skill your agent uses when writing production code test-first: the RED-GREEN-REFACTOR cycle, false-RED detection, vertical slicing, scope escalation, test process discipline, and code generation…

    164 GitHub stars~2.7k tokensUpdated today
    Auto-check: notes
  • Diff Driven Docs

    romiluz13/cc10x

    A skill your agent uses when a BUILD phase completes, a commit is staged, or a PR is about to be created, and the diff has not yet been reflected in documentation.

    164 GitHub stars~2.7k tokensUpdated today
    Auto-check: notes
  • Planning

    romiluz13/cc10x

    A skill your agent uses when writing an execution plan or a decision RFC: task decomposition, context references, validation levels, risk-based testing, ADR format, plan completeness gate, and…

    164 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Verification

    romiluz13/cc10x

    A skill your agent uses when judging whether a task reached its goal, not just finished: the gate function, self-critique gate, validation levels, evidence array protocol, and goal-backward lens.

    164 GitHub stars~1.9k tokensUpdated today
    Auto-check: notes
  • Agent Common

    romiluz13/cc10x

    A skill your agent uses when a cc10x agent starts a task: the shared preamble for the memory protocol, the contract format, and the output rules.

    164 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Cc10x Guide

    romiluz13/cc10x

    Answers questions about cc10x itself — what it is, how to install and configure it, how the router, workflows, memory, and hooks operate, and how to troubleshoot.

    164 GitHub stars~2.2k tokensUpdated today
    Auto-check passed

Categories

Questions about QA Strategy

What does QA Strategy do?

A skill your agent uses when designing or checking a QA test system: choosing the right tier, scenario matrices with pipeline-wide observation points, test-environment topology and isolation…. QA Strategy is an agent skill from romiluz13/cc10x. Use when designing or checking a QA test system: choosing the right tier, scenario matrices with pipeline-wide observation points, test-environment topology and isolation, fixture lifecycle, and flake sources.

When should I use QA Strategy?

QA Strategy fits situations like: checking a QA test system: choosing the right tier; scenario matrices with pipeline-wide observation points; test-environment topology and isolation; fixture lifecycle.

How do I install QA Strategy in Claude Code?

Run `npx skills add romiluz13/cc10x --skill qa-strategy -a claude-code`. Or copy the skill folder (plugins/cc10x/skills/qa-strategy in romiluz13/cc10x) into .claude/skills/qa-strategy in your project. Claude Code loads it when a task matches its description.

How do I install QA Strategy in Codex?

Run `npx skills add romiluz13/cc10x --skill qa-strategy -a codex`. Or copy the skill folder (plugins/cc10x/skills/qa-strategy in romiluz13/cc10x) into .agents/skills/qa-strategy in your project. Codex loads it when a task matches its description.

Can I use QA Strategy in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add romiluz13/cc10x --skill qa-strategy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/qa-strategy, .gemini/skills/qa-strategy, .github/skills/qa-strategy and .opencode/skills/qa-strategy in your project.

What does QA Strategy need to run?

Going by SKILL.md and its folder, QA Strategy needs the command-line tools its instructions call (docker). Our summary lists: Docker. Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash, LSP.

Does QA Strategy access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is QA Strategy safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does QA Strategy use?

QA Strategy is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does QA Strategy use?

About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to QA Strategy?

Skills that share tags, products or a category with QA Strategy: Testing OpenLogi UI (AprilNEA/OpenLogi, 23k stars), Testing Hashql (hashintel/hash, 1.7k stars), Dynamo Unit Testing (DynamoDS/Dynamo, 2k stars) and Designing Tests (CloudAI-X/opencode-workflow, 275 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains QA Strategy?

romiluz13 (a GitHub user) maintains it in romiluz13/cc10x, which has 164 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 7, 2026.

Source: romiluz13/cc10x on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.