Agent skill

Comprehensive Testing

by majiayu000 in majiayu000/spellbook

Complete testing strategy covering TDD workflow, test pyramid, unit/integration/E2E/property testing, framework best practices (Jest, Vitest, pytest), mock strategies, and CI integration.

MITAuto-check: notesTesting & QA

Install Comprehensive Testing

skills CLI
$ npx skills add majiayu000/spellbook --skill comprehensive-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install majiayu000/spellbook comprehensive-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/majiayu000/spellbook.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/comprehensive-testing .claude/skills/comprehensive-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
comprehensive-testing
GitHub stars
286
Token cost
~2.6k tokens
SKILL.md length
230 words
Files
2
Skills in repo
97
Repo updated
First seen
Licence
MIT

At a glance

Complete testing strategy covering TDD workflow, test pyramid, unit/integration/E2E/property testing, framework best practices (Jest, Vitest, pytest), mock strategies, and CI integration.

  • Works in 6 steps: Write Tests First → Verify Tests Fail → Commit Test Suite → …
  • Reviewing test quality
  • SKILL.md covers Core Philosophy, Test Pyramid, TDD Workflow (Anthropic… and Test Structure Patterns, plus 3 more sections
  • Calls git and npm

What it does

Comprehensive Testing is an agent skill from majiayu000/spellbook. Complete testing strategy covering TDD workflow, test pyramid, unit/integration/E2E/property testing, framework best practices (Jest, Vitest, pytest), mock strategies, and CI integration. Use when writing tests, reviewing test quality, or establishing testing standards.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `reference/extended.md`).

It sits in Testing & QA, covering Unit testing and Test strategy. It works with pytest, Vitest and Jest. The repository describes itself as: Cross-runtime skills for Claude Code, Codex, and multi-agent workflows. The licence is MIT.

When your agent uses it

  • Reviewing test quality
  • Establishing testing standards

Example prompts

  • “/comprehensive-testing”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Bash, Edit, Write

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Write Tests First
  2. Verify Tests Fail
  3. Commit Test Suite
  4. Implement Incrementally
  5. Verify with Subagent
  6. Commit Implementation

What it can do on your machine

Read from SKILL.md and the folder at commit 6310f81. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Bash
    • Edit
    • Write

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • anthropic.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Comprehensive Testing loads about 2.6k tokens when it runs. Until then it costs about 73 tokens; SKILL.md has 230 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~73
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Grep, Glob, Bash, Edit, Write

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from majiayu000/spellbook at commit 6310f81, republished under its MIT licence (© majiayu000). 230 words, ~2,615 tokens.

Download SKILL.mdSave it as .claude/skills/comprehensive-testing/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
comprehensive-testing
description
Complete testing strategy covering TDD workflow, test pyramid, unit/integration/E2E/property testing, framework best practices (Jest, Vitest, pytest), mock strategies, and CI integration. Use when writing tests, reviewing test quality, or establishing testing standards.
allowed-tools
Read, Grep, Glob, Bash, Edit, Write

Comprehensive Testing

Based on Anthropic's Claude Code Best Practices and community patterns

Core Philosophy

"Claude performs best when it has a clear target to iterate against—a test case provides concrete success criteria."

Testing is not about proving code works; it's about designing code that is testable and documenting expected behavior.


Test Pyramid

        /\
       /  \        E2E Tests (10%)
      /----\       - Full user flows
     /      \      - Slowest, most brittle
    /--------\
   /          \    Integration Tests (20%)
  /------------\   - Component interaction
 /              \  - Real dependencies
/----------------\
       Unit Tests (70%)
       - Single function/method
       - Fast, isolated, many
LevelSpeedScopeWhen to Use
Unit<10msSingle functionAll logic
Integration<1sMultiple componentsAPIs, DB
E2E<30sFull flowCritical paths

The 6-Step Process
1. WRITE TESTS FIRST
   ↓
2. VERIFY TESTS FAIL
   ↓
3. COMMIT TEST SUITE
   ↓
4. IMPLEMENT CODE
   ↓
5. VERIFY WITH SUBAGENT
   ↓
6. COMMIT IMPLEMENTATION
Step 1: Write Tests First
markdown
Be EXPLICIT about TDD to avoid mock implementations:

"I want to implement [feature] using TDD.
First, write tests for [expected behavior] with these input/output pairs:
- Input: X → Expected: Y
- Input: A → Expected: B
Do NOT create any implementation yet."
Step 2: Verify Tests Fail
bash
# Run tests and confirm they fail for the RIGHT reason
npm test  # or pytest, go test, etc.

# Expected: "function not found" or "undefined"
# NOT: syntax error, wrong import
Step 3: Commit Test Suite
bash
git add tests/
git commit -m "test: Add tests for [feature] (RED phase)"
Step 4: Implement Incrementally
markdown
"Now implement the code to make these tests pass.
Do NOT modify the tests.
Run tests after each change until all pass."

Claude will enter an autonomous loop:

Write code → Run tests → Analyze failures → Adjust → Repeat
Step 5: Verify with Subagent
markdown
"Use a subagent to independently verify the implementation:
- Is it overfitting to tests?
- Are edge cases handled?
- Is the code maintainable?"
Step 6: Commit Implementation
bash
git add src/
git commit -m "feat: Implement [feature] (GREEN phase)"

Test Structure Patterns

AAA Pattern (Arrange-Act-Assert)
typescript
describe('UserService', () => {
  it('should create user with valid email', async () => {
    // Arrange - Setup test data and dependencies
    const userRepo = new InMemoryUserRepository();
    const service = new UserService(userRepo);
    const input = { email: 'test@example.com', name: 'Test' };

    // Act - Execute the code under test
    const user = await service.create(input);

    // Assert - Verify the results
    expect(user.email).toBe('test@example.com');
    expect(user.id).toBeDefined();
    expect(await userRepo.findById(user.id)).toEqual(user);
  });
});
Given-When-Then Pattern
python
def test_order_total_with_discount():
    """
    Given an order with items totaling $100
    When a 20% discount is applied
    Then the total should be $80
    """
    # Given
    order = Order()
    order.add_item(Item(price=50))
    order.add_item(Item(price=50))

    # When
    order.apply_discount(Percentage(20))

    # Then
    assert order.total == Money(80)

Framework Best Practices

Jest / Vitest (JavaScript/TypeScript)
typescript
// Structure
describe('ModuleName', () => {
  describe('methodName', () => {
    it('should [expected behavior] when [condition]', () => {});
  });
});

// Setup/Teardown
beforeAll(async () => { /* one-time setup */ });
beforeEach(() => { /* per-test setup */ });
afterEach(() => { /* per-test cleanup */ });
afterAll(async () => { /* one-time cleanup */ });

// Async testing
it('handles async operations', async () => {
  const result = await asyncFunction();
  expect(result).toBe(expected);
});

// Error testing
it('throws on invalid input', () => {
  expect(() => validate(null)).toThrow('Input required');
});

// Snapshot testing (use sparingly)
it('renders correctly', () => {
  const tree = renderer.create(<Component />).toJSON();
  expect(tree).toMatchSnapshot();
});

// Table-driven tests
it.each([
  [1, 1, 2],
  [2, 2, 4],
  [0, 0, 0],
])('add(%i, %i) = %i', (a, b, expected) => {
  expect(add(a, b)).toBe(expected);
});

vitest.config.ts:

typescript
export default defineConfig({
  test: {
    globals: true,
    environment: 'node',
    coverage: {
      provider: 'v8',
      reporter: ['text', 'json', 'html'],
      thresholds: {
        lines: 80,
        branches: 70,
        functions: 80,
      },
    },
  },
});
pytest (Python)
python
import pytest
from mymodule import Calculator

# Fixtures for dependency injection
@pytest.fixture
def calculator():
    return Calculator()

@pytest.fixture
def database():
    db = TestDatabase()
    yield db
    db.cleanup()

# Parametrized tests
@pytest.mark.parametrize("a,b,expected", [
    (1, 1, 2),
    (2, 2, 4),
    (0, 0, 0),
    (-1, 1, 0),
])
def test_add(calculator, a, b, expected):
    assert calculator.add(a, b) == expected

# Exception testing
def test_divide_by_zero(calculator):
    with pytest.raises(ZeroDivisionError):
        calculator.divide(1, 0)

# Async testing
@pytest.mark.asyncio
async def test_async_operation():
    result = await async_function()
    assert result == expected

# Markers for categorization
@pytest.mark.slow
@pytest.mark.integration
def test_database_connection(database):
    assert database.is_connected()

conftest.py:

python
import pytest

@pytest.fixture(scope="session")
def database_url():
    return "postgresql://test:test@localhost/test"

@pytest.fixture(autouse=True)
def reset_database(database):
    yield
    database.rollback()

pytest.ini:

ini
[pytest]
testpaths = tests
python_files = test_*.py
python_functions = test_*
addopts = -v --cov=src --cov-report=term-missing
markers =
    slow: marks tests as slow
    integration: marks tests as integration tests
Go Testing
go
package mypackage

import (
    "testing"
    "github.com/stretchr/testify/assert"
    "github.com/stretchr/testify/require"
)

func TestAdd(t *testing.T) {
    tests := []struct {
        name     string
        a, b     int
        expected int
    }{
        {"positive numbers", 1, 2, 3},
        {"zero values", 0, 0, 0},
        {"negative numbers", -1, -2, -3},
    }

    for _, tt := range tests {
        t.Run(tt.name, func(t *testing.T) {
            result := Add(tt.a, tt.b)
            assert.Equal(t, tt.expected, result)
        })
    }
}

// Table-driven with subtests
func TestUserService_Create(t *testing.T) {
    t.Run("creates user with valid input", func(t *testing.T) {
        repo := NewInMemoryRepo()
        svc := NewUserService(repo)

        user, err := svc.Create(CreateUserInput{Email: "test@example.com"})

        require.NoError(t, err)
        assert.NotEmpty(t, user.ID)
        assert.Equal(t, "test@example.com", user.Email)
    })

    t.Run("returns error for invalid email", func(t *testing.T) {
        repo := NewInMemoryRepo()
        svc := NewUserService(repo)

        _, err := svc.Create(CreateUserInput{Email: "invalid"})

        require.Error(t, err)
        assert.Contains(t, err.Error(), "invalid email")
    })
}

Mock Strategy

When to Mock
ScenarioMock?Reason
External APIs✅ YesSlow, unreliable, costs money
Time/Date✅ YesNon-deterministic
Random✅ YesNon-deterministic
Database (unit)✅ YesSlow, complex setup
Database (integration)❌ NoTest real behavior
Your own code⚠️ RarelyPrefer real implementations
File system⚠️ DependsUse temp dirs when possible
How to Mock
typescript
// Jest - Mock module
jest.mock('./emailService', () => ({
  sendEmail: jest.fn().mockResolvedValue({ success: true }),
}));

// Jest - Mock function
const mockCallback = jest.fn();
mockCallback.mockReturnValue(42);

// Vitest - Spy
import { vi } from 'vitest';
const spy = vi.spyOn(console, 'log');

// Time mocking
beforeEach(() => {
  vi.useFakeTimers();
  vi.setSystemTime(new Date('2024-01-01'));
});

afterEach(() => {
  vi.useRealTimers();
});
python
# pytest - Mock with unittest.mock
from unittest.mock import Mock, patch, MagicMock

@patch('mymodule.external_api.fetch')
def test_with_mocked_api(mock_fetch):
    mock_fetch.return_value = {'data': 'mocked'}
    result = my_function()
    assert result == expected
    mock_fetch.assert_called_once_with('expected_arg')

# Fixture-based mock
@pytest.fixture
def mock_email_service():
    service = Mock()
    service.send.return_value = True
    return service
Prefer Test Doubles Over Mocks
typescript
// ❌ Heavy mocking
const mockRepo = {
  findById: jest.fn().mockResolvedValue({ id: '1', name: 'Test' }),
  save: jest.fn().mockResolvedValue(undefined),
  delete: jest.fn().mockResolvedValue(undefined),
};

// ✅ In-memory implementation (test double)
class InMemoryUserRepository implements UserRepository {
  private users: Map<string, User> = new Map();

  async findById(id: string): Promise<User | null> {
    return this.users.get(id) || null;
  }

  async save(user: User): Promise<User> {
    this.users.set(user.id, user);
    return user;
  }

  async delete(id: string): Promise<void> {
    this.users.delete(id);
  }
}

Extended Reference

Detailed material starting at ## Boundary Testing has been moved to reference/extended.md to keep this skill concise. Load that reference when the task requires the moved examples, command catalogs, checklists, platform details, or implementation templates.

© majiayu000, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/comprehensive-testing of majiayu000/spellbook.

  • SKILL.md
  • reference/extended.md

Open the folder on GitHubat commit 6310f81

Compare with similar skills

Comprehensive Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Comprehensive Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Comprehensive Testing this skillmajiayu000/spellbook286—~2.6kAutomated safety check: NotesMIT
Designing TestsCloudAI-X/claude-workflow-v21.4k1 repos~1.5kAutomated safety check: PassMIT
Test GuardamElnagdy/guard-skills1.3k2 repos~2.1kAutomated safety check: PassMIT
Agent Harness Testing Methodologyhuiliyi37/Tianshu-harness1.1k—~1kAutomated safety check: NotesApache-2.0
MoAI TDD Workflowmodu-ai/moai-adk1.2k—~3.1kAutomated safety check: PassApache-2.0
Test-Driven Development Enforcerzereight/gitlab-mcp2k1 repos~904Automated safety check: PassMIT

Similar skills

  • Designing Tests

    CloudAI-X/claude-workflow-v2

    Designs and implements testing strategies for any codebase. An agent skill from CloudAI-X/claude-workflow-v2.

    1.4k GitHub starsUsed in 1 repo~1.5k tokens
    Testing & QAAuto-check passed
  • Test Guard

    amElnagdy/guard-skills

    Reviews newly written or edited tests against nine rules that cut test bloat, such as mock-heavy checks and near-duplicate cases, before they are committed.

    1.3k GitHub starsUsed in 2 repos~2.1k tokens
    Testing & QAAuto-check passed
  • Agent Harness Testing Methodology

    huiliyi37/Tianshu-harness

    Guides an agent through probing an unfamiliar project's test setup, then choosing a red-light-first testing strategy matched to the task type.

    1.1k GitHub stars~1k tokensUpdated today
    Testing & QAAuto-check: notes
  • MoAI TDD Workflow

    modu-ai/moai-adk

    Drives test-first development through the RED, GREEN, REFACTOR cycle, with a config switch that selects between TDD and a DDD workflow for existing code.

    1.2k GitHub stars~3.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Enforces strict red-green-refactor, with a failing test first, the minimum code to pass it, then cleanup, and a quick reference for common test runners.

    2k GitHub starsUsed in 1 repo~904 tokens
    Testing & QAAuto-check passed
  • Python Testing Patterns

    jh941213/my-cc-harness

    Implement comprehensive testing strategies with pytest, fixtures, mocking, and test-driven development.

    126 GitHub starsUsed in 17 repos~5.4k tokens
    Testing & QAAuto-check passed

More from majiayu000/spellbook

All 97 skills in this repo
  • Skill Ecosystem Doctor

    majiayu000/spellbook

    Audits and repairs how coding-agent Skills are owned, copied and exposed across runtimes, from canonical sources to quarantine and retirement.

    286 GitHub stars~3k tokensUpdated yesterday
    Auto-check passed
  • AGENTS.md Scaffold

    majiayu000/spellbook

    Scans a repository for real evidence and proposes, or on request writes, a small stack of root and scoped AGENTS.md files with validation commands and generated-file boundaries.

    286 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Product Demo Builder

    majiayu000/spellbook

    Plans, produces or diagnoses evidence-backed product demo videos: script, capture plan, pacing checks and verified final media built on real product behavior.

    286 GitHub stars~3.3k tokensUpdated yesterday
    Auto-check passed
  • Flowguard Task Guard

    majiayu000/spellbook

    Single entry point that routes long or ambiguous agent tasks, checks live state, bounds autonomous loops and leaves a resumable handoff.

    286 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • npm Supply Chain Check

    majiayu000/spellbook

    Scans a repository, its lockfiles and node_modules for known malicious npm package versions and install-time indicators, using a read-only Python scanner.

    286 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Product Manager Toolkit

    majiayu000/spellbook

    Product management helpers: a RICE scoring script, an interview transcript analyzer and PRD templates for prioritizing features, synthesizing research and writing requirements.

    286 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Comprehensive Testing

What does Comprehensive Testing do?

Complete testing strategy covering TDD workflow, test pyramid, unit/integration/E2E/property testing, framework best practices (Jest, Vitest, pytest), mock strategies, and CI integration. Comprehensive Testing is an agent skill from majiayu000/spellbook. Complete testing strategy covering TDD workflow, test pyramid, unit/integration/E2E/property testing, framework best practices (Jest, Vitest, pytest), mock strategies, and CI integration.

When should I use Comprehensive Testing?

Comprehensive Testing fits situations like: reviewing test quality; establishing testing standards.

How do I install Comprehensive Testing in Claude Code?

Run `npx skills add majiayu000/spellbook --skill comprehensive-testing -a claude-code`. Or copy the skill folder (skills/comprehensive-testing in majiayu000/spellbook) into .claude/skills/comprehensive-testing in your project. Claude Code loads it when a task matches its description.

How do I install Comprehensive Testing in Codex?

Run `npx skills add majiayu000/spellbook --skill comprehensive-testing -a codex`. Or copy the skill folder (skills/comprehensive-testing in majiayu000/spellbook) into .agents/skills/comprehensive-testing in your project. Codex loads it when a task matches its description.

Can I use Comprehensive Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/spellbook --skill comprehensive-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/comprehensive-testing, .gemini/skills/comprehensive-testing, .github/skills/comprehensive-testing and .opencode/skills/comprehensive-testing in your project.

What does Comprehensive Testing need to run?

Going by SKILL.md and its folder, Comprehensive Testing needs the command-line tools its instructions call (git and npm). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash, Edit, Write.

Does Comprehensive Testing access the network?

SKILL.md names 1 domain. As links in the text: anthropic.com. This is read from the text; nothing was executed.

Is Comprehensive Testing safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Comprehensive Testing use?

Comprehensive Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Comprehensive Testing use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Comprehensive Testing?

Skills that share tags, products or a category with Comprehensive Testing: Designing Tests (CloudAI-X/claude-workflow-v2, 1.4k stars), Test Guard (amElnagdy/guard-skills, 1.3k stars), Agent Harness Testing Methodology (huiliyi37/Tianshu-harness, 1.1k stars) and MoAI TDD Workflow (modu-ai/moai-adk, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Comprehensive Testing?

majiayu000 (a GitHub user) maintains it in majiayu000/spellbook, which has 286 GitHub stars. The repository holds 97 skills in this directory. The repository was last updated on October 8, 2026.

Source: majiayu000/spellbook on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.