Agent skill

Testing Strategy

by AnastasiyaW in AnastasiyaW/codex-claude-code-config

A skill your agent uses when planning or reviewing tests for a code change, choosing between unit, focused regression, integration, contract, end-to-end, performance, security, property-based, or…

MITAuto-check passedTesting & QA

Install Testing Strategy

skills CLI
$ npx skills add AnastasiyaW/codex-claude-code-config --skill testing-strategy -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AnastasiyaW/codex-claude-code-config testing-strategy --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AnastasiyaW/codex-claude-code-config.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/development/testing-strategy .claude/skills/testing-strategy && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
testing-strategy
GitHub stars
154
Token cost
~2k tokens
SKILL.md length
1,049 words
Files
1
Skills in repo
50
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when planning or reviewing tests for a code change, choosing between unit, focused regression, integration, contract, end-to-end, performance, security, property-based, or…

  • Works in 9 steps: Freeze the acceptance criteria as… → Inspect the changed files and classify… → Select the lowest useful test level from… → …
  • Reviewing tests for a code change
  • SKILL.md covers Workflow, Compact Matrix, Test Kinds and Agent Evidence Contract, plus 3 more sections
  • Calls git

What it does

Testing Strategy is an agent skill from AnastasiyaW/codex-claude-code-config. Use when planning or reviewing tests for a code change, choosing between unit, focused regression, integration, contract, end-to-end, performance, security, property-based, or agent-evaluation checks; classify change risk first and select the smallest evidence set that proves the behavior. Do not use for a single obvious test command, pure documentation changes, or a full security audit without a testing question.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Test strategy, Agent evaluation and testing and Security review. The repository describes itself as: Claude Code, Codex, and multi-agent configuration system: principles, hooks, skills, and workflow patterns for AI-assisted development. The licence is MIT.

When your agent uses it

  • Reviewing tests for a code change
  • Choosing between unit
  • Focused regression
  • Agent-evaluation checks

Example prompts

  • “/testing-strategy”

Workflow steps

9 steps, taken from the first numbered list in SKILL.md.

  1. Freeze the acceptance criteria as observable outcomes.
  2. Inspect the changed files and classify the risk.
  3. Select the lowest useful test level from the matrix below and name the profile.
  4. Run the fast gate first. If it fails, fix the cause before adding more tests.
  5. Add a focused regression test for a confirmed bug or a changed invariant.
  6. Test real boundaries only when the change crosses them.
  7. Keep security-proof and release-attestation checks out of staging-smoke unless
  8. For high-risk or long-horizon work, use a fresh-context verifier and store
  9. When a verified stage becomes the input to another stage, seal that boundary

What it can do on your machine

Read from SKILL.md and the folder at commit 67709af. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Testing Strategy loads about 2k tokens when it runs. Until then it costs about 109 tokens; SKILL.md has 1,049 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~109
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AnastasiyaW/codex-claude-code-config at commit 67709af, republished under its MIT licence (© AnastasiyaW). 1,049 words, ~2,016 tokens.

Download SKILL.mdSave it as .claude/skills/testing-strategy/SKILL.md (or your agent's skills folder).
name
testing-strategy
description
Use when planning or reviewing tests for a code change, choosing between unit, focused regression, integration, contract, end-to-end, performance, security, property-based, or agent-evaluation checks; classify change risk first and select the smallest evidence set that proves the behavior. Do not use for a single obvious test command, pure documentation changes, or a full security audit without a testing question.

Testing Strategy

Testing is an evidence-selection problem, not a contest to run the largest suite. Choose the smallest set that can falsify the changed behavior, then add one higher-level check only when it covers a boundary the lower level cannot. Keep execution environments reusable, but separate their evidence profiles: staging-smoke, security-proof, release-attestation, and nightly-stress. Use harness-feedback when a gate is reported as overloaded or misplaced.

Workflow

  1. Freeze the acceptance criteria as observable outcomes.
  2. Inspect the changed files and classify the risk.
  3. Select the lowest useful test level from the matrix below and name the profile.
  4. Run the fast gate first. If it fails, fix the cause before adding more tests.
  5. Add a focused regression test for a confirmed bug or a changed invariant.
  6. Test real boundaries only when the change crosses them.
  7. Keep security-proof and release-attestation checks out of staging-smoke unless the acceptance criteria explicitly require that evidence.
  8. For high-risk or long-horizon work, use a fresh-context verifier and store the command, revision, result, and skipped checks in a durable artifact.
  9. When a verified stage becomes the input to another stage, seal that boundary with commit/tree, contract, input/output digests, and a fresh verdict. Mark an unavailable external prerequisite as BLOCKED; do not rerun unrelated accepted code merely because the following environment is unavailable.

Compact Matrix

ChangeRequired evidenceUsually deferred
Docs, comments, formatting onlyLink/lint check when relevantRuntime suite
Pure function, local refactorFast checks + focused unit/regression testsFull E2E, mutation
Parser, serializer, file, DB, API adapterFast + focused + one real boundary/integration checkBrowser E2E unless user flow changes
Auth, permissions, migrations, concurrency, public API, deploymentFast + focused + integration/contract + targeted smoke; independent review for non-trivial changesFull load test unless performance is in scope
UI or user journeyFast + component/focused checks + one stable E2E smokeLarge browser matrix
Release or performance claimAll applicable lower levels + fixed benchmark/security/release evidenceNothing that is part of the claim

The Stop hook runs the project's fast/default suite only when the working tree contains code or test changes. Projects with a complex suite may declare .claude/test-policy.json:

json
{
  "fast": ["python", "-m", "pytest", "-q", "tests/unit"],
  "integration": ["python", "-m", "pytest", "-q", "tests/integration"],
  "release": ["python", "-m", "pytest", "-q"]
}

fast is the automatic Stop gate. integration is additionally selected for high-risk changes when present. release is explicit or CI-only; do not make every edit pay the release-suite cost. Optional profiles make the separation explicit; a staging profile must not contain release-signing requirements.

Test Kinds

  • Unit: isolated behavior and invariants; fast and numerous.
  • Focused regression: a minimal test that was red before a fix and green after it. Keep it when it protects a real contract.
  • Integration: one real boundary such as a database, filesystem, queue, or external adapter. Use a local/test dependency, never production.
  • Contract: provider/consumer schema and serialization expectations.
  • Smoke/E2E: a small number of real user or release paths; keep them stable.
  • Property/fuzz: invariants over generated inputs; use for parsers, normalizers, state machines, and edge-heavy algorithms.
  • Performance/security: only when the change or release claim needs it; preserve a fixed workload and baseline.
  • Agent eval: test task completion, tool selection, recovery, and safety on a versioned golden set. Deterministic assertions come first; an LLM judge is an additional signal, never the sole proof of code correctness.

Agent Evidence Contract

An agent must report: revision, changed scope, commands actually run, exit status, relevant counts, environment constraints, and checks not run with a reason. “Tests passed” without command output or a durable evidence file is not proof. A generated test is a candidate until it reproduces the failure or asserts a stable contract; do not add broad snapshot tests merely to inflate coverage.

For a confirmed bug use bug-reproducer: reproduce first, then fix, then run the same test again. For a large/high-risk change use proof-verify: a fresh context must produce the final verdict. For a safe structural refactor use refactoring-safely and characterization tests before the transformation.

Show full SKILL.md (413 more words)Show less

Anti-Duplication Rules

  • If a higher-level test finds a failure with no lower-level failure, add the smallest lower-level reproducer and keep the higher-level test only if it proves a distinct boundary.
  • Do not run unit, integration, E2E, benchmark, and security suites by default just because they exist. Route by changed boundary and risk.
  • Do not use retries, sleeps, snapshots, or skip/xfail to make red tests look green. A flaky test needs a cause, a bounded quarantine reason, or a fix.
  • Do not claim release readiness from a fast suite alone.
  • Do not treat a new external blocker as a failed upstream candidate. Verify the stage that changed; promote the same sealed input when the prerequisite arrives.

Gotchas

  • A green test suite proves only the exercised behavior; it does not prove absence of defects.
  • Mocks can hide serialization and wiring failures. Keep one real boundary test for each important adapter.
  • End-to-end tests are valuable but expensive and flaky; they should protect journeys, not duplicate every branch already tested below.
  • Mutation testing is a periodic test-quality audit, not a per-edit gate. On Windows, verify the tool's runtime requirements before adding it to CI.
  • Agent trajectories need task-outcome checks and tool-call checks, not only final-text similarity.
  • A VM harness is an execution environment, not a release profile. Reuse the VM for staging and security checks, but attach signing and artifact identity checks only to release-attestation.

Troubleshooting

SymptomLikely causeAction
Stop gate runs in a docs-only changeProject has no Git-visible status or a broad overrideCheck git status; keep the default command scoped in .claude/test-policy.json
Fast suite passes, integration failsA real boundary was changed or mocked awayAdd/fix the boundary test; do not weaken the fast gate
E2E is flakyTiming, shared state, browser/environment dependencyMake state isolated and waits explicit; reduce E2E to a stable smoke
Generated test passes without exposing the bugTest asserts implementation details or never goes redReproduce the pre-fix failure and assert the user-visible invariant
Agent claims completion with skipped checksMissing evidence contract or verifierRecord the skip reason and run proof-verify for high-risk work
Agent says the harness is too strict or blocks smokeProfiles are coupled or a gate is misplacedInvoke harness-feedback; capture the blocker, split profiles, and rerun the reduced smoke
A later audit asks to repeat a green earlier stageProof identity was not recorded, or its source/input changedCheck the stage receipt; reuse a sealed matching receipt or record a superseding stage

© AnastasiyaW, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/development/testing-strategy of AnastasiyaW/codex-claude-code-config.

Open the folder on GitHubat commit 67709af

Compare with similar skills

Testing Strategy next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Testing Strategy compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Testing Strategy this skillAnastasiyaW/codex-claude-code-config154—~2kAutomated safety check: PassMIT
Security UpdatesSmilyOrg/photofield608—~808Automated safety check: PassMIT
Moai Ref Testing Pyramidmodu-ai/moai-adk1.2k—~1.8kAutomated safety check: PassApache-2.0
Testing QAaiskillstore/marketplace4333 repos~1.2kAutomated safety check: PassNone
Testingnotque/vexjoy-agent441—~4.7kAutomated safety check: NotesMIT
Plan Testsgenkovich/sdd171—~3.1kAutomated safety check: PassMIT

Similar skills

  • Security Updates

    SmilyOrg/photofield

    Check and apply security updates across the photofield project (api/Go, ui/npm, docs/npm, e2e/npm).

    608 GitHub stars~808 tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Moai Ref Testing Pyramid

    modu-ai/moai-adk

    Test pyramid strategy, coverage targets, test patterns, and quality metrics reference.

    1.2k GitHub stars~1.8k tokensUpdated today
    Testing & QAAuto-check passed
  • Testing QA

    aiskillstore/marketplace

    Comprehensive testing and QA workflow covering unit testing, integration testing, E2E testing, browser automation, and quality assurance.

    433 GitHub starsUsed in 3 repos~1.2k tokens
    Testing & QAAuto-check passed
  • Testing

    notque/vexjoy-agent

    Testing: TDD, E2E, preferred patterns, test-value audits, verification, agent testing.

    441 GitHub stars~4.7k tokensUpdated today
    Testing & QAAuto-check: notes
  • Plan Tests

    genkovich/sdd

    A skill your agent uses to turn a feature's acceptance criteria into a test plan before any test is written — a table that maps every spec.md §5 acceptance criterion to at least one test, names the…

    171 GitHub stars~3.1k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.

    16k GitHub starsUsed in 1 repo~2.1k tokens
    Testing & QAAuto-check: notes

More from AnastasiyaW/codex-claude-code-config

All 50 skills in this repo
  • Bug Reproducer

    AnastasiyaW/codex-claude-code-config

    Find likely software bugs in a codebase, rank concrete bug candidates, and prove or reject them with focused regression tests before proposing a fix.

    154 GitHub stars~4.1k tokensUpdated today
    Auto-check passed
  • Motion Framer

    AnastasiyaW/codex-claude-code-config

    A skill your agent uses when implementing Motion or Framer Motion in React/JavaScript: interactive UI components, micro-interactions, gestures, layout or page transitions, and scroll-based animation.

    154 GitHub starsUsed in 1 repo~5.2k tokens
    Auto-check passed
  • Proof Verify

    AnastasiyaW/codex-claude-code-config

    Plan-based verification - freeze acceptance criteria before building, then verify after with an independent fresh-context agent (the builder must not verify their own work).

    154 GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Workflow Orchestration

    AnastasiyaW/codex-claude-code-config

    Написание и запуск Claude Code dynamic workflows (JS-оркестратор субагентов).

    154 GitHub stars~3.8k tokensUpdated today
    Auto-check passed
  • Notebooklm Grounded Research

    AnastasiyaW/codex-claude-code-config

    A skill your agent uses when: NotebookLM, notebooklm MCP, large documentation sets, courses, books, papers, or citation-backed research are mentioned.

    154 GitHub stars~2.4k tokensUpdated today
    Auto-check: warnings
  • Deepseek Provider Contract

    AnastasiyaW/codex-claude-code-config

    Validate a proposed DeepSeek API integration before any key or project context is sent: check thinking-mode tool-call history, strict-schema assumptions, bounded output, and provider data boundaries.

    154 GitHub stars~1.2k tokensUpdated today
    Auto-check passed

Questions about Testing Strategy

What does Testing Strategy do?

A skill your agent uses when planning or reviewing tests for a code change, choosing between unit, focused regression, integration, contract, end-to-end, performance, security, property-based, or…. Testing Strategy is an agent skill from AnastasiyaW/codex-claude-code-config. Use when planning or reviewing tests for a code change, choosing between unit, focused regression, integration, contract, end-to-end, performance, security, property-based, or agent-evaluation checks; classify change risk first and select the smallest evidence set that proves the behavior.

When should I use Testing Strategy?

Testing Strategy fits situations like: reviewing tests for a code change; choosing between unit; focused regression; agent-evaluation checks.

How do I install Testing Strategy in Claude Code?

Run `npx skills add AnastasiyaW/codex-claude-code-config --skill testing-strategy -a claude-code`. Or copy the skill folder (skills/development/testing-strategy in AnastasiyaW/codex-claude-code-config) into .claude/skills/testing-strategy in your project. Claude Code loads it when a task matches its description.

How do I install Testing Strategy in Codex?

Run `npx skills add AnastasiyaW/codex-claude-code-config --skill testing-strategy -a codex`. Or copy the skill folder (skills/development/testing-strategy in AnastasiyaW/codex-claude-code-config) into .agents/skills/testing-strategy in your project. Codex loads it when a task matches its description.

Can I use Testing Strategy in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AnastasiyaW/codex-claude-code-config --skill testing-strategy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/testing-strategy, .gemini/skills/testing-strategy, .github/skills/testing-strategy and .opencode/skills/testing-strategy in your project.

What does Testing Strategy need to run?

Going by SKILL.md and its folder, Testing Strategy needs the command-line tools its instructions call (git).

Does Testing Strategy access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Testing Strategy safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Testing Strategy use?

Testing Strategy is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Testing Strategy use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Testing Strategy?

Skills that share tags, products or a category with Testing Strategy: Security Updates (SmilyOrg/photofield, 608 stars), Moai Ref Testing Pyramid (modu-ai/moai-adk, 1.2k stars), Testing QA (aiskillstore/marketplace, 433 stars) and Testing (notque/vexjoy-agent, 441 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Testing Strategy?

AnastasiyaW (a GitHub user) maintains it in AnastasiyaW/codex-claude-code-config, which has 154 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on October 9, 2026.

Source: AnastasiyaW/codex-claude-code-config on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.