Agent skill

Test Strategy Hardening

by swyxio in swyxio/skills

Diagnose or improve tests as a system, including flaky or nondeterministic suites, duplicated coverage, expensive fixtures, slow feedback, and test selection.

MITAuto-check passedTesting & QA

Install Test Strategy Hardening

skills CLI
$ npx skills add swyxio/skills --skill test-strategy-hardening -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install swyxio/skills test-strategy-hardening --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/test-strategy-hardening .claude/skills/test-strategy-hardening && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-strategy-hardening
GitHub stars
176
Token cost
~1.1k tokens
SKILL.md length
529 words
Files
2
Skills in repo
89
Repo updated
First seen
Licence
MIT

At a glance

Diagnose or improve tests as a system, including flaky or nondeterministic suites, duplicated coverage, expensive fixtures, slow feedback, and test selection.

  • Works in 4 steps: Name the behavior that could regress and… → Search for existing coverage of the same… → Choose the lowest boundary that can… → …
  • Explicitly makes test architecture
  • SKILL.md covers Counterweight: confidence, not…, Define the contract, Choose a faithful boundary and Remove nondeterminism, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Test Strategy Hardening is an agent skill from swyxio/skills. Diagnose or improve tests as a system, including flaky or nondeterministic suites, duplicated coverage, expensive fixtures, slow feedback, and test selection. Use only when the user explicitly makes test architecture or suite behavior the primary task. Do not trigger for ordinary implementation work, adding or fixing a regression test, resolving one product bug, or running existing checks.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in Testing & QA, covering Test strategy. The repository describes itself as: Agent skills for Claude Code and other AI agents. The licence is MIT.

When your agent uses it

  • Explicitly makes test architecture
  • Suite behavior the primary task
  • Ordinary implementation work
  • Fixing a regression test

Example prompts

  • “/test-strategy-hardening”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Name the behavior that could regress and who observes it.
  2. Search for existing coverage of the same invariant.
  3. Choose the lowest boundary that can prove the behavior faithfully.
  4. Identify time, randomness, concurrency, storage, network, global-state, and teardown dependencies.

What it can do on your machine

Read from SKILL.md and the folder at commit 038ef34. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Strategy Hardening loads about 1.1k tokens when it runs. Until then it costs about 104 tokens; SKILL.md has 529 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~104
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from swyxio/skills at commit 038ef34, republished under its MIT licence (© swyxio). 529 words, ~1,076 tokens.

Download SKILL.mdSave it as .claude/skills/test-strategy-hardening/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
test-strategy-hardening
description
Diagnose or improve tests as a system, including flaky or nondeterministic suites, duplicated coverage, expensive fixtures, slow feedback, and test selection. Use only when the user explicitly makes test architecture or suite behavior the primary task. Do not trigger for ordinary implementation work, adding or fixing a regression test, resolving one product bug, or running existing checks.

Test Strategy Hardening

Improve confidence per minute without inflating test count or ceremony.

Counterweight: confidence, not test inventory

  • Start from the decision the suite must support and the plausible defect it must catch.
  • Search for replacement coverage before adding a test. Delete or consolidate tests when confidence remains equivalent.
  • Do not create unit, integration, end-to-end, snapshot, fixture, and contract layers as a completeness ladder.
  • Prefer fixing product nondeterminism or simplifying the seam over building elaborate test harnesses around it.
  • Treat coverage percentage, test count, suite breadth, and green full-suite runs as signals, not goals.
  • Do not refactor unrelated tests, rewrite fixtures, add retries, expand CI matrices, or run global suites solely because test code was touched.
  • Preserve focused failure localization; deduplication is not valuable when it makes diagnosis harder.
  • Stop when the named suite problem is resolved and proportionally verified. Report remaining blind spots without filling them speculatively.

Define the contract

Before changing a test suite:

  1. Name the behavior that could regress and who observes it.
  2. Search for existing coverage of the same invariant.
  3. Choose the lowest boundary that can prove the behavior faithfully.
  4. Identify time, randomness, concurrency, storage, network, global-state, and teardown dependencies.

Do not add a test merely because a source line changed. A useful test must fail for a plausible defect and pass for the intended behavior.

Choose a faithful boundary

  • Use pure tests for transformations, validation, policies, and invariants.
  • Use component or module tests for framework behavior.
  • Use integration tests when storage, protocols, processes, or concurrency define the contract.
  • Reserve end-to-end tests for critical user journeys and runtime ownership proofs.

Do not mock away the behavior under test. Prefer stable, externally meaningful assertions over private calls, generated values, broad snapshots, or expectations copied from production logic.

Show full SKILL.md (235 more words)Show less

Remove nondeterminism

  • Inject or freeze time.
  • Seed or inject randomness and ID generation.
  • Isolate mutable state, environment variables, ports, providers, and caches.
  • Await background work and close resources in teardown.
  • Synchronize through observable events or bounded polling against a real condition.

Larger timeouts, retries, global serialization, skipped tests, and weaker assertions are not flake fixes. Reproduce the smallest failure, classify the cause, fix the source, and repeat the repaired test enough to exercise the former failure mode.

Consolidate deliberately

  • Use table-driven cases when one contract varies only by input and outcome.
  • Extract repeated realistic setup into feature-owned fixtures.
  • Keep scenario-specific decisions visible at the call site.
  • Retain focused tests when they materially improve failure localization.
  • Delete or merge a test only after naming the replacement coverage.

Measure setup, imports, execution, and teardown before optimizing. Share immutable baseline construction only when mutable state remains isolated and invalidation is complete.

Validate proportionally

During implementation, expand validation with risk:

  1. the changed or nearest focused test;
  2. related tests and affected package checks;
  3. cross-package contracts when a shared boundary changes;
  4. the full suite only when repository policy or a shared high-risk boundary requires it.

Never report a focused suite as proof that the full suite passed.

Report

Summarize:

  • contracts strengthened;
  • tests merged or removed and their replacement coverage;
  • flaky-test causes and deterministic fixes;
  • commands actually run and suites intentionally omitted;
  • runtime changes and remaining blind spots.

© swyxio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in test-strategy-hardening of swyxio/skills.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 038ef34

Compare with similar skills

Test Strategy Hardening next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Strategy Hardening compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Strategy Hardening this skillswyxio/skills176—~1.1kAutomated safety check: PassMIT
Testing OpenLogi UIAprilNEA/OpenLogi23k—~1.1kAutomated safety check: PassApache-2.0
Testing Hashqlhashintel/hash1.7k—~1.9kAutomated safety check: PassAGPL-3.0
Dynamo Unit TestingDynamoDS/Dynamo2k—~622Automated safety check: PassApache-2.0
Bmad Testarch Test Designbmad-code-org/bmad-method-test-architecture-enterprise1053 repos~1.4kAutomated safety check: PassCustom licence
Test RoadmapOvid/paad131—~3.5kAutomated safety check: PassMIT

Similar skills

  • Testing OpenLogi UI

    AprilNEA/OpenLogi

    Verifies OpenLogi's native GPUI interface with focused tests, the component gallery and a mock agent, choosing the evidence that fits each change.

    23k GitHub stars~1.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Testing Hashql

    hashintel/hash

    HashQL testing strategies including compiletest (UI tests), unit tests, and snapshot tests.

    1.7k GitHub stars~1.9k tokensUpdated today
    Testing & QAAuto-check passed
  • Dynamo Unit Testing

    DynamoDS/Dynamo

    Write comprehensive NUnit tests for the Dynamo codebase following Dynamo testing patterns, conventions, and architectural constraints.

    2k GitHub stars~622 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Bmad Testarch Test Design

    bmad-code-org/bmad-method-test-architecture-enterprise

    Create system-level or epic-level test plans. An agent skill from bmad-code-org/bmad-method-test-architecture-enterprise.

    105 GitHub starsUsed in 3 repos~1.4k tokens
    Testing & QAAuto-check passed
  • Test Roadmap

    Ovid/paad

    Analyzes a repository and any existing test suite, grades existing tests for weakness, classifies mocks, emits a phased roadmap for building a test suite that catches real regressions, then executes…

    131 GitHub stars~3.5k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Designing Tests

    CloudAI-X/opencode-workflow

    Guides test strategy, TDD/BDD approaches, test coverage planning, and testing best practices.

    275 GitHub stars~2.9k tokensUpdated 9 mo ago
    Testing & QAAuto-check passed

More from swyxio/skills

All 89 skills in this repo
  • Programmatic Agents

    swyxio/skills

    Run a selected coding-agent CLI programmatically, with latency, error, usage, cost, and trace logging.

    176 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Design, implement, audit, or refresh protected username and handle namespaces for public products.

    176 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • New Mac Setup

    swyxio/skills

    Fully automated new Mac setup for fullstack web developers and AI engineers.

    176 GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • Youtube API

    swyxio/skills

    Manage YouTube videos programmatically via the YouTube Data API v3 — upload video files, upload custom thumbnails, update video metadata (titles, descriptions, tags), and query video/channel info…

    176 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Batch YouTube Studio upload workflow for videos sourced from Airtable, Google Drive, Loom, YouTube, or local files.

    176 GitHub stars~1.5k tokensUpdated today
    Auto-check: warnings
  • Reconstruct and visually analyze paired agent, game, or policy trajectories to determine whether changed actions produced their intended effects.

    176 GitHub stars~1.8k tokensUpdated today
    Auto-check passed

Categories

Questions about Test Strategy Hardening

What does Test Strategy Hardening do?

Diagnose or improve tests as a system, including flaky or nondeterministic suites, duplicated coverage, expensive fixtures, slow feedback, and test selection. Test Strategy Hardening is an agent skill from swyxio/skills. Diagnose or improve tests as a system, including flaky or nondeterministic suites, duplicated coverage, expensive fixtures, slow feedback, and test selection.

When should I use Test Strategy Hardening?

Test Strategy Hardening fits situations like: explicitly makes test architecture; suite behavior the primary task; ordinary implementation work; fixing a regression test.

How do I install Test Strategy Hardening in Claude Code?

Run `npx skills add swyxio/skills --skill test-strategy-hardening -a claude-code`. Or copy the skill folder (test-strategy-hardening in swyxio/skills) into .claude/skills/test-strategy-hardening in your project. Claude Code loads it when a task matches its description.

How do I install Test Strategy Hardening in Codex?

Run `npx skills add swyxio/skills --skill test-strategy-hardening -a codex`. Or copy the skill folder (test-strategy-hardening in swyxio/skills) into .agents/skills/test-strategy-hardening in your project. Codex loads it when a task matches its description.

Can I use Test Strategy Hardening in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add swyxio/skills --skill test-strategy-hardening -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-strategy-hardening, .gemini/skills/test-strategy-hardening, .github/skills/test-strategy-hardening and .opencode/skills/test-strategy-hardening in your project.

What does Test Strategy Hardening need to run?

SKILL.md names no scripts, command-line tools or credentials: Test Strategy Hardening is instructions for the agent only.

Does Test Strategy Hardening access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Test Strategy Hardening safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Strategy Hardening use?

Test Strategy Hardening is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Strategy Hardening use?

About 1.1k tokens (SKILL.md is roughly 4.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test Strategy Hardening?

Skills that share tags, products or a category with Test Strategy Hardening: Testing OpenLogi UI (AprilNEA/OpenLogi, 23k stars), Testing Hashql (hashintel/hash, 1.7k stars), Dynamo Unit Testing (DynamoDS/Dynamo, 2k stars) and Bmad Testarch Test Design (bmad-code-org/bmad-method-test-architecture-enterprise, 105 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Strategy Hardening?

swyxio (a GitHub user) maintains it in swyxio/skills, which has 176 GitHub stars. The repository holds 89 skills in this directory. The repository was last updated on October 5, 2026.

Source: swyxio/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.