Agent skill

Test Architect

by FerroxLabs in FerroxLabs/wayland

Test strategy designer covering test pyramid, coverage strategies, fixture design, mock patterns, TDD/BDD workflows, property-based testing, and test isolation.

Apache-2.0Auto-check passedTesting & QA

Install Test Architect

skills CLI
$ npx skills add FerroxLabs/wayland --skill test-architect -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install FerroxLabs/wayland test-architect --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/software-engineering/test-architect .claude/skills/test-architect && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-architect
GitHub stars
608
Token cost
~3.5k tokens
SKILL.md length
1,111 words
Files
1
Skills in repo
1,194
Repo updated
First seen
Licence
Apache-2.0

At a glance

Test strategy designer covering test pyramid, coverage strategies, fixture design, mock patterns, TDD/BDD workflows, property-based testing, and test isolation.

  • Works in 4 steps: Every test creates its own data. Never… → Use builders/factories, not raw… → Use realistic but deterministic data. No… → …
  • The user asks about test architect
  • SKILL.md covers Test Pyramid, Testing Trophy (Alternative…, Test Categorization and Test Naming Conventions, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Test Architect is an agent skill from FerroxLabs/wayland. Test strategy designer covering test pyramid, coverage strategies, fixture design, mock patterns, TDD/BDD workflows, property-based testing, and test isolation. Use when the user asks about test architect, test architect best practices, or needs guidance on test architect implementation. Do NOT use when the user needs a different specialized skill or is asking about an unrelated technology domain.

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Test strategy and Test-driven development. The repository describes itself as: Wayland - The AI Agent That Perceives. Reasons. Acts. Evolves. The licence is Apache-2.0.

When your agent uses it

  • The user asks about test architect
  • Test architect best practices
  • Needs guidance on test architect implementation
  • The user needs a different specialized skill

Example prompts

  • “/test-architect”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Every test creates its own data. Never share mutable state across tests.
  2. Use builders/factories, not raw constructors. Tests should not break when a field is added.
  3. Use realistic but deterministic data. No random() in fixtures unless testing randomness.
  4. Clean up after tests that create external state (database rows, files, queue messages).

What it can do on your machine

Read from SKILL.md and the folder at commit 4c030c7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python, javascript, java, typescript, gherkin and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Architect loads about 3.5k tokens when it runs. Until then it costs about 104 tokens; SKILL.md has 1,111 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~104
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from FerroxLabs/wayland at commit 4c030c7, republished under its Apache-2.0 licence (© FerroxLabs). 1,111 words, ~3,456 tokens.

Download SKILL.mdSave it as .claude/skills/test-architect/SKILL.md (or your agent's skills folder).
name
test-architect
description
Test strategy designer covering test pyramid, coverage strategies, fixture design, mock patterns, TDD/BDD workflows, property-based testing, and test isolation. Use when the user asks about test architect, test architect best practices, or needs guidance on test architect implementation. Do NOT use when the user needs a different specialized skill or is asking about an unrelated technology domain.
license
Apache-2.0
metadata.author
foundry-skills
metadata.version
1.0.0
metadata.tags
best-practices testing guide
metadata.category
software-engineering
metadata.subcategory
developer-tools
metadata.disclaimer
none
metadata.difficulty
intermediate

Test Architect

You are an expert test architect. Design test strategies that maximize defect detection while minimizing maintenance cost. Write tests that document intent, catch regressions, and enable fearless refactoring.

Test Pyramid

The test pyramid is a cost-optimization framework. Tests at the bottom are fast, cheap, and stable. Tests at the top are slow, expensive, and brittle.

         /  E2E  \           Few (5-10%)    Slow, brittle, high confidence
        /----------\
       / Integration \       Some (15-25%)  Moderate speed, moderate confidence
      /----------------\
     /    Unit Tests     \   Many (60-80%)  Fast, stable, low confidence per test
    /____________________\
Unit Tests (60-80% of tests)
  • Test a single function/method in isolation.
  • No I/O, no database, no network, no filesystem.
  • Execute in milliseconds.
  • Use mocks/stubs for external dependencies.
Integration Tests (15-25% of tests)
  • Test the interaction between 2+ components.
  • May use real database (in-memory or containerized).
  • Verify serialization, query correctness, API contracts.
  • Execute in seconds.
End-to-End Tests (5-10% of tests)
  • Test complete user workflows through the real system.
  • Use browser automation (Playwright, Cypress) or API calls.
  • Cover critical business paths only (login, checkout, payment).
  • Execute in minutes.

Testing Trophy (Alternative Model)

Kent C. Dodds' testing trophy emphasizes integration tests:

         /  E2E  \           Few
        /----------\
       / Integration \       MOST (focus here)
      /----------------\
     /      Unit        \    Some
    /____________________\
   /    Static Analysis   \  Linting, types
  /________________________\

Use this model when your application is primarily about composing components (React apps, API services).

Test Categorization

By Scope
CategoryWhat it testsDependenciesSpeed
UnitSingle function/classNone (mocked)< 10ms
IntegrationComponent interactionReal DB, queues< 5s
ContractAPI compatibilitySchema/types< 1s
E2EFull user workflowEntire system< 60s
SmokeSystem is aliveDeployed system< 30s
By Purpose
CategoryGoalWhen to run
RegressionCatch broken behaviorEvery commit
AcceptanceVerify requirementsBefore release
PerformanceDetect speed regressionsNightly
SecurityFind vulnerabilitiesBefore release
ChaosTest resiliencePeriodically

Test Naming Conventions

Use descriptive names that form readable sentences.

Pattern: methodName_scenario_expectedBehavior
java
calculateTotal_emptyCart_returnsZero()
calculateTotal_singleItem_returnsItemPrice()
calculateTotal_withDiscount_appliesPercentOff()
Pattern: should [expected] when [condition]
javascript
it('should return zero when cart is empty')
it('should apply discount when coupon is valid')
it('should throw error when payment fails')
Pattern: given [context] when [action] then [outcome]
python
def test_given_valid_credentials_when_login_then_returns_token():
def test_given_expired_token_when_access_api_then_returns_401():

Fixture Design

Builder Pattern for Test Data
typescript
class UserBuilder {
  private user: Partial<User> = {
    name: "Test User",
    email: "test@example.com",
    role: "viewer",
    active: true,
  };

  withName(name: string): this { this.user.name = name; return this; }
  withRole(role: Role): this { this.user.role = role; return this; }
  asAdmin(): this { this.user.role = "admin"; return this; }
  inactive(): this { this.user.active = false; return this; }
  build(): User { return { ...this.user } as User; }
}

// Usage
const admin = new UserBuilder().asAdmin().build();
const inactiveUser = new UserBuilder().inactive().build();
Factory Functions
python
def make_order(
    status="pending",
    item_count=1,
    total=100.00,
    customer=None,
):
    customer = customer or make_customer()
    items = [make_item() for _ in range(item_count)]
    return Order(status=status, items=items, total=total, customer=customer)
Fixture Rules
  1. Every test creates its own data. Never share mutable state across tests.
  2. Use builders/factories, not raw constructors. Tests should not break when a field is added.
  3. Use realistic but deterministic data. No random() in fixtures unless testing randomness.
  4. Clean up after tests that create external state (database rows, files, queue messages).

Mock vs Stub vs Spy

Stub

Returns canned responses. Does not verify calls.

python
# Stub: always returns a fixed exchange rate
currency_service = Mock()
currency_service.get_rate.return_value = 1.25
Mock

Returns canned responses AND verifies interactions.

python
# Mock: verify that the email was sent with correct args
email_service = Mock()
process_order(order, email_service)
email_service.send.assert_called_once_with(
    to=order.customer.email,
    subject="Order Confirmation"
)
Spy

Wraps the real implementation, records calls.

javascript
const spy = jest.spyOn(analytics, 'track');
submitForm(data);
expect(spy).toHaveBeenCalledWith('form_submitted', data);
spy.mockRestore();
When to Use Each
TechniqueUse whenAvoid when
Real dependencyFast, deterministic, no side effectsSlow, flaky, or has side effects
StubNeed to control return valuesNeed to verify interaction
MockNeed to verify interaction occurredOver-specifying implementation details
SpyNeed real behavior + call recordingWant full isolation
FakeNeed simplified working implementationReal implementation is simple enough
Mock Antipatterns
  • Over-mocking: Mocking so much that the test verifies wiring, not behavior.
  • Mock-as-assertion: Using mock verification as the only assertion (fragile).
  • Mocking what you own: Mock boundaries, not your own classes. If you mock your own code, tests break when you refactor.

Property-Based Testing

Instead of testing with specific examples, test with properties that hold for all inputs.

Example
python
from hypothesis import given, strategies as st

@given(st.lists(st.integers()))
def test_sort_preserves_length(xs):
    assert len(sorted(xs)) == len(xs)

@given(st.lists(st.integers()))
def test_sort_is_idempotent(xs):
    assert sorted(sorted(xs)) == sorted(xs)

@given(st.lists(st.integers()))
def test_sort_is_ordered(xs):
    result = sorted(xs)
    for i in range(len(result) - 1):
        assert result[i] <= result[i + 1]
Common Properties
  • Round-trip: decode(encode(x)) == x
  • Idempotent: f(f(x)) == f(x)
  • Commutativity: f(a, b) == f(b, a)
  • Invariant preservation: len(f(xs)) == len(xs)
  • Oracle: Compare a fast implementation against a known-correct slow one
When to Use Property-Based Testing
  • Parsing (encode/decode)
  • Serialization
  • Data transformations
  • Mathematical operations
  • Sorting/searching algorithms

TDD Workflow

Red-Green-Refactor Cycle
  1. RED: Write a test that fails (the test compiles but the assertion fails).
  2. GREEN: Write the minimum code to make the test pass.
  3. REFACTOR: Clean up the code while keeping tests green.
TDD Rules (Robert C. Martin)
  1. Write no production code except to make a failing test pass.
  2. Write only enough of a test to demonstrate a failure.
  3. Write only enough production code to make the test pass.
Test List Technique

Before coding, write a list of all test cases you can think of:

- empty list returns empty
- single element returns that element
- two elements in order stay in order
- two elements reversed get swapped
- duplicate elements are preserved
- negative numbers sort correctly
- already sorted list is unchanged

Work through the list one test at a time.

BDD Workflow

Gherkin Syntax
gherkin
Feature: User login

Scenario: Successful login with valid credentials
  Given a registered user with email "user@test.com"
  And the user has password "secure123"
  When the user logs in with email "user@test.com" and password "secure123"
  Then the response status should be 200
  And the response should contain an auth token
  And the auth token should expire in 24 hours

Scenario: Failed login with wrong password
  Given a registered user with email "user@test.com"
  When the user logs in with email "user@test.com" and password "wrong"
  Then the response status should be 401
  And the response should contain error "Invalid credentials"

Test Isolation Patterns

Database Isolation
  1. Transaction rollback: Wrap each test in a transaction, rollback after.
  2. Database per test: Create a fresh database for each test (slow but safe).
  3. Truncation: Truncate all tables between tests.
  4. Schema migration: Run migrations before the test suite, seed per test.
Show full SKILL.md (433 more words)Show less
Time Isolation
python
# Inject clock dependency, never use datetime.now() directly
class OrderService:
    def __init__(self, clock=datetime.now):
        self.clock = clock

    def is_expired(self, order):
        return self.clock() > order.expires_at

# In tests
fixed_time = datetime(2024, 1, 15, 12, 0, 0)
service = OrderService(clock=lambda: fixed_time)
Network Isolation
  • Use WireMock, MSW, or VCR to record/replay HTTP interactions.
  • Never make real HTTP calls in unit tests.
  • In integration tests, use containers (Testcontainers) for real services.
Filesystem Isolation
  • Use temp directories that are cleaned up after each test.
  • Use in-memory filesystems when available.
  • Never write to the project source tree in tests.

Coverage Strategy

What to Measure
  • Line coverage: Percentage of lines executed. Minimum useful metric.
  • Branch coverage: Percentage of if/else branches taken. More meaningful.
  • Mutation coverage: Percentage of code mutations caught by tests. Most meaningful, most expensive.
Coverage Targets
  • Do NOT mandate 100% coverage. It leads to low-value tests.
  • Target 80% line coverage as a baseline.
  • Target 100% coverage on critical business logic (payments, auth, data integrity).
  • Ignore generated code, configuration, and trivial getters/setters.
Coverage as a Ratchet
  • Never let coverage decrease. Use CI to enforce this.
  • Require coverage for new code, not retroactively for old code.

Test Maintenance

Signs of Unhealthy Tests
  • Tests break when you refactor without changing behavior.
  • Tests take > 10 minutes to run.
  • Tests fail intermittently (flaky).
  • Nobody trusts the test suite ("oh, that test always fails").
  • Adding a feature requires modifying 20 tests.
Fixing Flaky Tests
  1. Identify the flaky test (track failure rates).
  2. Determine the cause: timing, ordering, shared state, network.
  3. Fix the root cause, do not add retries.
  4. If you cannot fix it, quarantine it (separate suite) and track it.

When to Use

Use this skill when:

  • Designing or implementing test architect solutions
  • Reviewing or improving existing test architect approaches
  • Making architectural or implementation decisions about test architect
  • Learning test architect patterns and best practices
  • Troubleshooting test architect-related issues

Do NOT use this skill when:

  • The question is about a fundamentally different technology domain
  • A more specific sibling skill covers the exact topic needed
  • The user needs a complete hands-on tutorial rather than expert guidance

Output Format

markdown
# Test Architect Analysis

## Context Assessment
[Situation summary and constraints]

## Recommended Approach
[Primary recommendation with rationale]

## Implementation Steps
1. [Step with specific details]
2. [Step with specific details]
3. [Step with specific details]

## Trade-offs and Considerations
- [Key trade-off 1]
- [Key trade-off 2]

## Next Steps
- [Immediate action item]
- [Follow-up action item]

Example

Input: "Help me implement test architect for a medium-scale production application"

Output: A structured analysis covering current state assessment, recommended test architect approach with specific patterns, implementation roadmap with milestones, and risk mitigation strategies tailored to the application scale and constraints.

Edge Cases

  • Legacy system integration: When test architect must coexist with legacy approaches, provide a gradual migration path rather than a complete rewrite
  • Scale mismatch: When the solution complexity exceeds the project scale, recommend a simpler approach and note when to revisit
  • Team skill gaps: When the team lacks experience with the recommended approach, include learning resources and simpler alternatives
  • Conflicting requirements: When constraints conflict (e.g., performance vs. maintainability), explicitly state the trade-off and recommend based on stated priorities

© FerroxLabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in src/process/resources/skills-library/bodies/skills/software-engineering/test-architect of FerroxLabs/wayland.

Open the folder on GitHubat commit 4c030c7

Compare with similar skills

Test Architect next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Architect compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Architect this skillFerroxLabs/wayland608—~3.5kAutomated safety check: PassApache-2.0
Golang Testingantoniopaya22/go-rest-template1729 repos~4.2kAutomated safety check: PassNone
Agent Harness Testing Methodologyhuiliyi37/Tianshu-harness1.1k—~1kAutomated safety check: NotesApache-2.0
Designing TestsCloudAI-X/opencode-workflow275—~2.9kAutomated safety check: PassMIT
Python Testingaffaan-m/ECC274k6 repos~4.7kAutomated safety check: PassMIT
Rust Testingkurealnum/dotfiles2906 repos~2.9kAutomated safety check: PassNone

Similar skills

  • Golang Testing

    antoniopaya22/go-rest-template

    Go testing patterns including table-driven tests, subtests, benchmarks, fuzzing, and test coverage.

    172 GitHub starsUsed in 9 repos~4.2k tokens
    Testing & QAAuto-check passed
  • Agent Harness Testing Methodology

    huiliyi37/Tianshu-harness

    Guides an agent through probing an unfamiliar project's test setup, then choosing a red-light-first testing strategy matched to the task type.

    1.1k GitHub stars~1k tokensUpdated 2 days ago
    Testing & QAAuto-check: notes
  • Designing Tests

    CloudAI-X/opencode-workflow

    Guides test strategy, TDD/BDD approaches, test coverage planning, and testing best practices.

    275 GitHub stars~2.9k tokensUpdated 9 mo ago
    Testing & QAAuto-check passed
  • Python Testing

    affaan-m/ECC

    Python testing strategies using pytest, TDD methodology, fixtures, mocking, parametrization, and coverage requirements.

    274k GitHub starsUsed in 6 repos~4.7k tokens
    Testing & QAAuto-check passed
  • Rust Testing

    kurealnum/dotfiles

    Rust testing patterns including unit tests, integration tests, async testing, property-based testing, mocking, and coverage.

    290 GitHub starsUsed in 6 repos~2.9k tokens
    Testing & QAAuto-check passed
  • Testing

    NiceEval/NiceEval

    Inspect, select, author, modify, or review NiceEval E2E and Unit test owners.

    156 GitHub stars~2.1k tokensUpdated yesterday
    Testing & QAAuto-check passed

More from FerroxLabs/wayland

All 1,194 skills in this repo
  • Star Office Helper

    FerroxLabs/wayland

    Install, start, connect, and troubleshoot visualization companion projects for Aion/OpenClaw, with Star-Office-UI as the default recommendation.

    608 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check: notes
  • Openclaw Setup

    FerroxLabs/wayland

    OpenClaw usage expert: Helps you install, deploy, configure, and use OpenClaw personal AI assistant.

    608 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Tvcontrol Setup

    FerroxLabs/wayland

    Set up TVControl end to end: install the connector, start TradingView Desktop with its control port open, load a watchlist export, add the indicators they use, and leave a working chart.

    608 GitHub stars~5.7k tokensUpdated yesterday
    Auto-check passed
  • Ab Testing Specialist

    FerroxLabs/wayland

    End-to-end guide for designing, running, and analyzing A/B tests including experiment design, statistical significance, sample size calculation, common pitfalls, and advanced testing patterns.

    608 GitHub stars~3.7k tokensUpdated yesterday
    Auto-check passed
  • Academic Writer

    FerroxLabs/wayland

    Complete academic writing guide covering thesis and dissertation structure, journal article format using IMRaD, literature review methodology, citation management, the peer review process, and…

    608 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check passed
  • Accessibility Auditor

    FerroxLabs/wayland

    Web accessibility expertise covering WCAG 2.2 conformance, audit methodology, ARIA patterns, keyboard navigation, screen reader testing, focus management, form accessibility, and automated vs manual…

    608 GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Test Architect

What does Test Architect do?

Test strategy designer covering test pyramid, coverage strategies, fixture design, mock patterns, TDD/BDD workflows, property-based testing, and test isolation. Test Architect is an agent skill from FerroxLabs/wayland. Test strategy designer covering test pyramid, coverage strategies, fixture design, mock patterns, TDD/BDD workflows, property-based testing, and test isolation.

When should I use Test Architect?

Test Architect fits situations like: the user asks about test architect; test architect best practices; needs guidance on test architect implementation; the user needs a different specialized skill.

How do I install Test Architect in Claude Code?

Run `npx skills add FerroxLabs/wayland --skill test-architect -a claude-code`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/software-engineering/test-architect in FerroxLabs/wayland) into .claude/skills/test-architect in your project. Claude Code loads it when a task matches its description.

How do I install Test Architect in Codex?

Run `npx skills add FerroxLabs/wayland --skill test-architect -a codex`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/software-engineering/test-architect in FerroxLabs/wayland) into .agents/skills/test-architect in your project. Codex loads it when a task matches its description.

Can I use Test Architect in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add FerroxLabs/wayland --skill test-architect -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-architect, .gemini/skills/test-architect, .github/skills/test-architect and .opencode/skills/test-architect in your project.

What does Test Architect need to run?

SKILL.md names no scripts, command-line tools or credentials: Test Architect is instructions for the agent only. Our summary lists: Python 3.

Does Test Architect access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Test Architect safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Architect use?

Test Architect is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Architect use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test Architect?

Skills that share tags, products or a category with Test Architect: Golang Testing (antoniopaya22/go-rest-template, 172 stars), Agent Harness Testing Methodology (huiliyi37/Tianshu-harness, 1.1k stars), Designing Tests (CloudAI-X/opencode-workflow, 275 stars) and Python Testing (affaan-m/ECC, 274k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Architect?

FerroxLabs (a GitHub user) maintains it in FerroxLabs/wayland, which has 608 GitHub stars. The repository holds 1,194 skills in this directory. The repository was last updated on October 6, 2026.

Source: FerroxLabs/wayland on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.